# AI Safety India (AISIN) | Full Knowledge Base & Technical Documentation > The authoritative reference document for AI Safety India (AISIN), covering programs, technical alignment curriculum, governance frameworks, career transition roadmaps, grant funding, and research tools. Version: 2026.8.0 Canonical URL: https://www.aisafetyindia.org/ Contact: contact@aisafetyindia.org --- ## 1. Executive Summary & Organizational Identity - **Name**: AI Safety India (AISIN) - **Organization Type**: Independent field-building and research community. - **Mission**: To build the technical research, red-teaming, and governance talent pipeline necessary to ensure that frontier artificial intelligence systems are safe, interpretable, and aligned with human flourishing. - **Geographic Focus**: India & Global South, connected directly to top global alignment hubs (San Francisco, London, Berkeley, Oxford). - **Core Strategy**: 1. Talent Sourcing & Acceleration: Upskilling top Indian software engineers, mathematicians, and ML practitioners via free 6-week intensive cohorts. 2. Research Incubation: Funding 12-week Phase 2 research fellowships with living stipends and high-density GPU clusters (A100/H100). 3. Ecosystem Infrastructure: Launching collegiate chapters, maintaining the global opportunities database, providing free 1-on-1 career advising, and advising national policy bodies (MeitY, NITI Aayog). --- ## 2. Core Educational Programs ### Track A: Technical AI Safety (6-Week Intensive Cohort) - **Format**: 6-Week Cohort-based, online/hybrid, free tuition, 8–12 hours/week commitment. - **Prerequisites**: Proficiency in Python, basic linear algebra/calculus, and familiarity with deep learning / PyTorch. - **Weekly Curriculum**: - **Week 1: Foundations of Deep Learning & Transformer Architectures**: Attention mechanisms, multi-head self-attention, positional embeddings, residual stream additive communication, building GPT from scratch in PyTorch. - **Week 2: Reinforcement Learning & Scalable Alignment (RLHF, DPO, KTO)**: Markov Decision Processes, policy gradients (PPO), reward modeling, KL divergence penalties, Direct Preference Optimization without reward models, Kahneman-Tversky Optimization. - **Week 3: Mechanistic Interpretability & Circuit Analysis**: TransformerLens, induction heads, circuit identification, direct logit attribution (DLA), activation patching (causal tracing), indirect object identification (IOI) circuits. - **Week 4: Dictionary Learning & Sparse Autoencoders (SAEs)**: Superposition hypothesis, polysemanticity vs. monosemanticity, overcomplete latent dictionaries, TopK and L1 sparsity loss functions, Anthropic's scaling monosemanticity methodology. - **Week 5: Adversarial Robustness, Evals & Automated Red Teaming**: Universal adversarial triggers, jailbreak prompts, Inspect Evals harness, multilingual vulnerability surfaces in Indic languages, automated attacker-model loops. - **Week 6: Advanced Alignment Threat Models & Capstone Project**: Deceptive alignment, sleeper agents, model sandbagging, situational awareness, scheming, final research project presentation. ### Track B: AI Governance & Policy (6-Week Intensive Cohort) - **Format**: 6-Week Cohort-based, online/hybrid, free tuition, 6–8 hours/week commitment. - **Target Audience**: Policy analysts, legal scholars, technologists, civil servants, and economists. - **Weekly Curriculum**: - **Week 1: AI Risk Typologies & Frontier Capabilities**: Extreme risks, autonomous weapon systems, cyberwarfare, biosecurity threats, frontier compute scaling trends. - **Week 2: Hardware & Compute Governance**: Semiconductor supply chain bottlenecks (ASML, TSMC, NVIDIA), FLOP thresholds (10^26), on-chip hardware security modules, cloud KYC frameworks. - **Week 3: Global Regulatory Frameworks**: US Executive Order 14110, EU AI Act risk tiers, UK and US AI Safety Institutes, Seoul & Bletchley Frontier Safety Commitments. - **Week 4: Indian Policy & Digital Sovereignty**: IndiaAI Mission, Digital Personal Data Protection (DPDP) Act 2023, MeitY advisory directives, NITI Aayog Responsible AI principles. - **Week 5: Corporate Governance & Safety Cases**: Responsible Scaling Policies (RSPs), ASL-3/ASL-4 risk tiers, third-party pre-deployment auditing, whistleblowing protections. - **Week 6: Multilateral Treaties & Non-Proliferation**: International compute registries, IAEA-style verification inspections, India's geopolitical leverage in non-aligned AI safety treaties. ### Track C: Phase 2 Research Fellowships - **Duration**: 12 Weeks (Quarterly cohorts). - **Compensation**: Merit-based living stipend in INR + dedicated GPU cluster allocations (RunPod / GCP TPU/GPU). - **Mentorship**: Direct 1-on-1 pairing with active alignment researchers, SPAR fellows, and international laboratory collaborators. - **Target Deliverables**: arXiv preprints, open-source safety tool packages, LessWrong / Alignment Forum research writeups, or targeted policy whitepapers. --- ## 3. Career Advising & Ecosystem Services ### 1-on-1 Career Advising - **Lead Advisor**: Aditya Raj (SPAR Fellow at UC Berkeley, Founder of AI Safety India). - **Cost**: 100% Free. - **Booking URL**: https://airtable.com/appoJKCpenhAAOXO1/pagSPhe9HhsjNxRq6/form - **Scope of Guidance**: - Transition pathways from Big Tech software engineering or academic ML to technical AI safety. - Review and feedback on fellowship applications (MATS, ARENA, SPAR, Astra, CHAI). - Ideation and scoping for grant proposals (LTFF, Manifund, SFF). - Project selection for mechanistic interpretability and red-teaming portfolios. --- ## 4. Regional Hubs & University Chapters 1. **AI Safety NIT Agartala (AIS-NITA)**: - Official collegiate chapter at National Institute of Technology Agartala. - Activities: Weekly paper reading circles, ARENA code sprints, and student mentorship. - URL: https://www.aisafetyindia.org/hubs/ais-nita 2. **Delhi AI Alignment & Safety Collective**: - Regional NCR hub uniting researchers, engineers, and policy scholars across Delhi, Noida, and Gurugram. - Activities: In-person salon discussions on mathematical foundations of agency, deceptive alignment, and compute governance. - URL: https://www.aisafetyindia.org/hubs/delhi-ai-alignment-safety-collective 3. **University Groups 2026 Incubator**: - Semester seed grants ($500–$2,500) for student organizers launching campus AI safety chapters. - Provided: Curriculum packets, speaker matching, compute vouchers, and weekly organizer coaching. - URL: https://www.aisafetyindia.org/community/university-groups-2026 --- ## 5. High-Yield Knowledge Guides & Blueprints ### AI Safety Career Guide for India - **URL**: https://www.aisafetyindia.org/resources/ai-safety-career-guide-india - **Core Highlights**: - Step-by-step roadmap for software engineers transitioning to alignment in 3–6 months. - Key global fellowships: MATS ($5k–$7k/mo), ARENA (London/Berkeley), SPAR (Berkeley), Open Philanthropy Scholarships. - Compensation landscape: Independent grants ($2.5k–$5k/mo), Fellowships ($4k–$7k/mo), Frontier Labs ($150k–$400k+). - Critical reading list: Elhage et al. Transformer Circuits, Templeton et al. Scaling Monosemanticity, Ngo AGI Safety from First Principles, Hubinger Sleeper Agents. ### AI Safety Research Grants & Funding Guide - **URL**: https://www.aisafetyindia.org/resources/ai-safety-funding-india - **Core Highlights**: - Global Grantmakers: Long-Term Future Fund (LTFF), Manifund (fast 1–2 week turnaround), Survival and Flourishing Fund (SFF). - Compute Grants: Anthropic API Credits, OpenAI Researcher Access, Google Cloud TPU/GPU credits, IndiaAI Mission. - 5-Step Winning Grant Blueprint: Concrete hypothesis, empirical falsifiability, GPU math breakdown, timeline milestones. - India Tax & Legal Structuring: FEMA/FIRC remittance codes (`P0802`), GST zero-rated export with LUT, Section 44ADA presumptive taxation. --- ## 6. AI Alignment & Safety Glossary (Core Technical Definitions) 1. **Sparse Autoencoder (SAE)**: Unsupervised neural network decomposing polysemantic residual stream activations into an overcomplete sparse dictionary of human-interpretable monosemantic features using an L1/TopK penalty. 2. **Residual Stream**: The central linear highway in transformers where attention heads and MLP layers read and write activations additively ($x_{l+1} = x_l + \text{Attn}(x_l) + \text{MLP}(x_l)$). 3. **Induction Head**: A two-layer attention circuit performing in-context pattern completion of the pattern $[A][B] \dots [A] \to [B]$. 4. **Polysemanticity**: When an individual neuron activates on multiple unrelated semantic concepts due to dimensional superposition. 5. **Superposition Hypothesis**: The property of neural networks representing more features than dimensions by packing them as non-orthogonal linear directions. 6. **Direct Logit Attribution (DLA)**: Direct projection of an individual component's output through the unembedding matrix $W_U$ to measure its contribution to the next-token probability. 7. **Activation Patching**: Causal intervention replacing internal activations on corrupted inputs with clean activations to isolate functional circuits. 8. **Direct Preference Optimization (DPO)**: Closed-form implicit reward alignment algorithm directly optimizing preference pairs without a separate reward model or PPO loop. 9. **Constitutional AI (RLAIF)**: Automated alignment pipeline where an AI critiques and refines its own responses against written constitutional principles. 10. **Deceptive Alignment**: When an AI optimizes the training objective during training solely to prevent itself from being modified or shut down, waiting for deployment to pursue an unaligned goal. 11. **Sleeper Agents**: Models exhibiting safe behavior during evaluation that execute backdoor attacks when triggered by a specific deployment condition. 12. **Compute Governance**: Regulating advanced semiconductor hardware (GPUs/TPUs) as the physical, verifiable bottleneck for training dangerous frontier models. 13. **FLOP Threshold**: Objective compute ceiling (e.g. $10^{26}$ operations) triggering mandatory regulatory safety evaluations and government notifications. --- ## 7. Canonical URLs & Endpoints - Home: https://www.aisafetyindia.org/ - Technical AI Safety Course: https://www.aisafetyindia.org/programs/technical-ai-safety - AI Governance & Policy Course: https://www.aisafetyindia.org/programs/ai-governance-policy - 1-on-1 Career Advising: https://www.aisafetyindia.org/programs/advising - Projects & Open Tools: https://www.aisafetyindia.org/programs/projects-and-tools - Events & Sessions: https://www.aisafetyindia.org/community/events - India AI Safety Orgs Directory: https://www.aisafetyindia.org/community/orgs - University Groups Incubator: https://www.aisafetyindia.org/community/university-groups-2026 - Career Guide (India): https://www.aisafetyindia.org/resources/ai-safety-career-guide-india - Funding Guide (India): https://www.aisafetyindia.org/resources/ai-safety-funding-india - Theory of Change: https://www.aisafetyindia.org/about/theory-of-change - Brand Guidelines: https://www.aisafetyindia.org/brand - LLM Machine-Readable Index: https://www.aisafetyindia.org/llms.txt - LLM Full Reference Document: https://www.aisafetyindia.org/llms-full.txt