Arnav Raj

Dual Degree, Bachelor's + Master's (B.Tech + M.Tech), Computer Science & Engineering, IIT Delhi

I'm in my final year of CS at IIT Delhi. I work on RLHF, mostly reward models and what they get wrong about the people whose preferences they are meant to encode. I think the averaging step, where thousands of annotators collapse into one number, throws away most of what made the feedback worth collecting. I've worked at Google DeepMind (IC via Topcoder), Abundant AI (YC F24), Georgia Tech, and Harvard's Edge Computing Lab. Four papers this year: two at ICML 2026 workshops, one at ICLR 2026 GRaM, and KG-MuLQA at ACL 2026 as a coauthor.

AI Safety LLM Evaluation Interpretability RLHF Reinforcement Learning
Arnav Raj portrait

Google DeepMind (IC via Topcoder LLC)

AI Evaluation & SFT
Jun 2026 to Oct 2026 · Mountain View, CA (Remote)
  • Worked on RSI.
  • Worked on Gemini evaluations.
  • Worked on Colab Bench and ML RL environments.

Abundant AI (YC F24)

Contractor
Nov 2025 to Jun 2026 · San Francisco, CA (Remote)
  • Built GPU-enabled long-running RL environments.
  • Built SWE and data-science RL environments.
  • Designed sandbox architectures.

Georgia Institute of Technology Financial Services Innovation Lab

Research Intern
May 2024 to Jun 2025 · Atlanta, GA (Remote)
  • Built the end-to-end pipeline behind KG-MuLQA (ACL 2026), generating 20,139 multi-hop QA pairs from knowledge graphs over 170 real financial documents.
  • Wrote the scoring harness the benchmark uses to evaluate frontier LLMs.

Harvard University Edge Computing Lab

Research Intern
May 2024 to Dec 2024 · Cambridge, MA (Remote)
  • Built a benchmarking framework for LLM-generated RTL hardware designs, with automated syntax checking, testbench verification, and power-performance-area analysis.
  • Compared prompting strategies across GPT-4 and Llama, re-prompting the designs that failed verification.

Reading Multi-Token Answers from a Single Hidden State

Arnav Raj
Under review, ISEAI

Solo-authored. Token lenses read a hidden state one token at a time, so they cannot recover an answer like "North Korea". Patching a late-layer state into an early layer of a copy prompt gets the model to spell the answer out itself, roughly doubling exact-match accuracy over patching at the layer the state came from.

PDF → Code →

PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration

Arnav Raj
Accepted at ICML 2026 Workshop on Pluralistic Alignment

Solo-authored. Introduced a lightweight statistical layer that calibrates RLHF reward models to each individual human annotator rather than collapsing thousands of raters into a single average. Improves alignment to genuine human preferences on standard pluralistic-alignment benchmarks while leaving the underlying reward model unchanged.

arXiv →

Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF

Arnav Raj
Accepted at ICML 2026 Workshop on RL from World Feedback (RLxF)

Solo-authored. Proposed a correction primitive for production RLHF pipelines where reward signals arrive late, such as slow code verifiers, large judge ensembles, and queued human review. Allows policy training to continue using delayed feedback instead of discarding it, with a closed-form proof of unbiasedness and substantial improvement over standard wait-or-drop strategies.

arXiv →

KG-MuLQA: Multi-hop Question Answering over Knowledge Graphs for Long-Context Evaluation

Nikita Tatarinov, B Vidhyakshaya Kannan, Haricharana Srinivasa, Arnav Raj, et al.
Accepted at ACL 2026

Benchmark for evaluating multi-hop reasoning in long-context language models, built from knowledge graphs over real-world financial documents. Designed and implemented the end-to-end pipeline for question generation, answer synthesis, and evaluation across frontier LLMs.

arXiv →

Hyperbolic Geometry of Reasoning: Probing LLM Hidden States

Arnav Raj
Accepted at ICLR 2026 Workshop on Geometry-grounded Representation Learning and Generative Modeling (GRaM)

Solo-authored. Interpretability study of how chain-of-thought reasoning models internally encode hierarchical structure. Showed that hyperbolic geometric probes recover this structure across the entire network while standard Euclidean probes fail in the deepest layers, evidence that reasoning models compress hierarchy into the representations responsible for the final answer.

OpenReview →

LLM Code Agent Evaluation Framework

Comprehensive evaluation suite for LLM code generation inspired by SWE-bench and MLE-bench, achieving 73% task completion with iterative self-correction.

Python Docker LangChain OpenAI API pytest AST

RL Agent for Code Optimization

PPO-based reinforcement learning agent that optimizes code performance through iterative refinement, achieving 35% runtime reduction and 20% memory improvement.

Python PyTorch OpenAI Gym Ray RLlib AST subprocess

Hangman AI: Transformer-Based Game Solver

Transformer-driven Hangman solver achieving >60% success using character-level modeling, morphological augmentation, and multi-strategy guess selection.

Python PyTorch NLTK Transformers NumPy
GitHub →

Graph Neural Network for User Personality Prediction

Bipartite user–product interaction graph leveraged with GNN architecture to infer user personality traits with high accuracy.

Python PyTorch PyTorch-Geometric
GitHub →

Data Search & Retrieval Engine

Neural-augmented retrieval system combining classical inverted indexing with dense vector search for hybrid relevance scoring across semi-structured technical documents.

Python Elasticsearch Faiss hnswlib FastAPI SentenceTransformers Docker
GitHub →

Context-Aware Spelling Correction

Noisy-channel + smoothed N-gram language modeling system achieving 88% accuracy on context-dependent spelling errors.

Python NLTK
GitHub →

TotalRecall: Cognitive Health App

Cross-platform Flutter app with AI-assisted memory support workflows for Alzheimer’s patients, backed by Firebase services.

Flutter Dart Firebase
GitHub →

Advanced Analytic Tool

A modular analytical platform that ingests heterogeneous datasets (CSV/JSON/SQL streams) and provides extensible pipelines for preprocessing, feature engineering, and rapid experimentation with classical ML and lightweight deep models.

Python Pandas Scikit-learn LightGBM FastAPI Docker
GitHub →

AI Player for Havannah Board Game

Monte Carlo Tree Search (MCTS) agent with UCB exploration achieving >80% win rate over strong RAVE-only baselines.

Python MCTS
GitHub →

SDN-based Intelligent Network Controller

High‑throughput OpenFlow controller with proactive L2 learning, shortest‑path routing, and loop prevention for complex topologies.

Python Ryu OpenFlow Mininet
GitHub →

Application-Layer Reliable Transport & Congestion Control

Implemented a reliable transport protocol over UDP plus Reno & CUBIC congestion control variants with detailed performance benchmarking.

Python UDP Mininet
GitHub →

OS Kernel Enhancements in xv6

Implemented a page swapping subsystem in the xv6 teaching OS, including victim selection, swap slot management, and page fault handling.

C x86 QEMU
GitHub →

I'm in my final year of CS at IIT Delhi, on the dual degree, bachelor's plus master's (B.Tech + M.Tech). I work on RLHF. The part I keep coming back to is the reward model: it gets trained on thousands of annotators who disagree with each other, and the standard recipe averages that disagreement away before training starts. I think that is a mistake, and most of what I have written so far is one or another attempt to show it.

Through 2026 I worked on AI evaluation and SFT for Google DeepMind, as an IC via Topcoder, on RSI, Gemini evaluations, and Colab Bench and ML RL environments. Before that I was at Abundant AI (YC F24), building GPU-enabled RL environments for long-running agents, SWE and data-science environments, and the sandbox architecture underneath them. Some adversarial ML benchmarks I wrote there ended up in the training pipelines of three of the world's top five AI labs, which I still find slightly absurd.

The newest one, under review at ISEAI, is about reading multi-token answers out of a single hidden state: token lenses give you one token at a time and cannot spell "North Korea", so instead I patch a late-layer state into an early layer of a copy prompt and let the model say the answer itself. PEBS, at an ICML 2026 workshop, calibrates a reward model to each individual rater instead of to their average. Retroactive Advantage Correction, at the ICML 2026 RLxF workshop, is a closed-form fix for the case where the reward arrives after the step it belongs to, so you can train on it rather than drop it. A third paper, at the ICLR 2026 GRaM workshop, probes reasoning models with hyperbolic geometry and finds hierarchy in the deepest layers, where Euclidean probes lose it. I'm also a coauthor on KG-MuLQA, a long-context multi-hop QA benchmark at ACL 2026.

Earlier I was a research intern at Georgia Tech's Financial Services Innovation Lab, and at Harvard's Edge Computing Lab. I co-founded the AI Safety Club at IIT Delhi, and went through BlueDot Impact's alignment course and the ARENA curriculum.

Download Resume (PDF) →

IIT Delhi

Indian Institute of Technology Delhi

Bachelor's + Master's (B.Tech + M.Tech) in Computer Science & Engineering
2022-2027

NK Security Scholar

Merit-based scholarship awarded to top 30 students at IIT Delhi for academic and technical excellence

Smart India Hackathon

National Top 5 Finalist in both 2023 and 2024 editions (India's largest student innovation competition)

JEE Advanced 2022

All India Rank 1,158 out of 1,000,000+ candidates (top 0.1%)

KVPY SX Fellowship 2021

National science fellowship awarded by the Government of India and IISc Bangalore

Codeforces Expert

1700+ rating in competitive programming

National Science Olympiads

Top 250 Astronomy, Top 300 Chemistry in India

Founding Member & Technical Lead / AI Safety Club, IIT Delhi

2025 – Present

Co-founded student research group on AI alignment, interpretability, and evaluation. Led reading groups on mechanistic interpretability (TransformerLens) and safety frameworks. Completed BlueDot Impact AI safety training and ARENA curriculum.

Technical Consultant / STEM AI Hackathon 2026 (AI-Collab Hack)

Jan 2026 – Present

Providing technical mentorship to 20+ teams building AI agents for STEM education at a hackathon jointly organized by IIT Delhi, Imperial College London, and Microsoft Garage.

Senior Editor / Tech Ambit (Pan-IIT Magazine)

2023 – 2025

Led 15-member editorial team across 23 IITs; curated and edited 30+ technical articles on AI and systems research.

Mess Secretary / Zanskar Hostel, IIT Delhi

Jun 2024 – May 2025

Elected by 400+ residents; managed operations team of 13. Awarded Best Mess Secretary for digitalization initiatives.

My research, mapped: each colour is a domain (safety & eval, RL & training, foundations & language); the bridges show where those domains meet. Click any topic for the papers behind it.