Arnav Raj
Dual Degree, Bachelor's + Master's (B.Tech + M.Tech), Computer Science & Engineering, IIT Delhi
I'm in my final year of CS at IIT Delhi. I work on RLHF, mostly reward models and what they get wrong about the people whose preferences they are meant to encode. I think the averaging step, where thousands of annotators collapse into one number, throws away most of what made the feedback worth collecting. I've worked at Google DeepMind (IC via Topcoder), Abundant AI (YC F24), Georgia Tech, and Harvard's Edge Computing Lab. Four papers this year: two at ICML 2026 workshops, one at ICLR 2026 GRaM, and KG-MuLQA at ACL 2026 as a coauthor.
Google DeepMind (IC via Topcoder LLC)
- Worked on RSI.
- Worked on Gemini evaluations.
- Worked on Colab Bench and ML RL environments.
Abundant AI (YC F24)
- Built GPU-enabled long-running RL environments.
- Built SWE and data-science RL environments.
- Designed sandbox architectures.
Georgia Institute of Technology Financial Services Innovation Lab
- Built the end-to-end pipeline behind KG-MuLQA (ACL 2026), generating 20,139 multi-hop QA pairs from knowledge graphs over 170 real financial documents.
- Wrote the scoring harness the benchmark uses to evaluate frontier LLMs.
Harvard University Edge Computing Lab
- Built a benchmarking framework for LLM-generated RTL hardware designs, with automated syntax checking, testbench verification, and power-performance-area analysis.
- Compared prompting strategies across GPT-4 and Llama, re-prompting the designs that failed verification.
Reading Multi-Token Answers from a Single Hidden State
Solo-authored. Token lenses read a hidden state one token at a time, so they cannot recover an answer like "North Korea". Patching a late-layer state into an early layer of a copy prompt gets the model to spell the answer out itself, roughly doubling exact-match accuracy over patching at the layer the state came from.
PDF → Code →PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration
Solo-authored. Introduced a lightweight statistical layer that calibrates RLHF reward models to each individual human annotator rather than collapsing thousands of raters into a single average. Improves alignment to genuine human preferences on standard pluralistic-alignment benchmarks while leaving the underlying reward model unchanged.
arXiv →Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF
Solo-authored. Proposed a correction primitive for production RLHF pipelines where reward signals arrive late, such as slow code verifiers, large judge ensembles, and queued human review. Allows policy training to continue using delayed feedback instead of discarding it, with a closed-form proof of unbiasedness and substantial improvement over standard wait-or-drop strategies.
arXiv →KG-MuLQA: Multi-hop Question Answering over Knowledge Graphs for Long-Context Evaluation
Benchmark for evaluating multi-hop reasoning in long-context language models, built from knowledge graphs over real-world financial documents. Designed and implemented the end-to-end pipeline for question generation, answer synthesis, and evaluation across frontier LLMs.
arXiv →Hyperbolic Geometry of Reasoning: Probing LLM Hidden States
Solo-authored. Interpretability study of how chain-of-thought reasoning models internally encode hierarchical structure. Showed that hyperbolic geometric probes recover this structure across the entire network while standard Euclidean probes fail in the deepest layers, evidence that reasoning models compress hierarchy into the representations responsible for the final answer.
OpenReview →LLM Code Agent Evaluation Framework
Comprehensive evaluation suite for LLM code generation inspired by SWE-bench and MLE-bench, achieving 73% task completion with iterative self-correction.
RL Agent for Code Optimization
PPO-based reinforcement learning agent that optimizes code performance through iterative refinement, achieving 35% runtime reduction and 20% memory improvement.
Hangman AI: Transformer-Based Game Solver
Transformer-driven Hangman solver achieving >60% success using character-level modeling, morphological augmentation, and multi-strategy guess selection.
Graph Neural Network for User Personality Prediction
Bipartite user–product interaction graph leveraged with GNN architecture to infer user personality traits with high accuracy.
Data Search & Retrieval Engine
Neural-augmented retrieval system combining classical inverted indexing with dense vector search for hybrid relevance scoring across semi-structured technical documents.
Context-Aware Spelling Correction
Noisy-channel + smoothed N-gram language modeling system achieving 88% accuracy on context-dependent spelling errors.
TotalRecall: Cognitive Health App
Cross-platform Flutter app with AI-assisted memory support workflows for Alzheimer’s patients, backed by Firebase services.
Advanced Analytic Tool
A modular analytical platform that ingests heterogeneous datasets (CSV/JSON/SQL streams) and provides extensible pipelines for preprocessing, feature engineering, and rapid experimentation with classical ML and lightweight deep models.
AI Player for Havannah Board Game
Monte Carlo Tree Search (MCTS) agent with UCB exploration achieving >80% win rate over strong RAVE-only baselines.
SDN-based Intelligent Network Controller
High‑throughput OpenFlow controller with proactive L2 learning, shortest‑path routing, and loop prevention for complex topologies.
Application-Layer Reliable Transport & Congestion Control
Implemented a reliable transport protocol over UDP plus Reno & CUBIC congestion control variants with detailed performance benchmarking.
OS Kernel Enhancements in xv6
Implemented a page swapping subsystem in the xv6 teaching OS, including victim selection, swap slot management, and page fault handling.
I'm in my final year of CS at IIT Delhi, on the dual degree, bachelor's plus master's (B.Tech + M.Tech). I work on RLHF. The part I keep coming back to is the reward model: it gets trained on thousands of annotators who disagree with each other, and the standard recipe averages that disagreement away before training starts. I think that is a mistake, and most of what I have written so far is one or another attempt to show it.
Through 2026 I worked on AI evaluation and SFT for Google DeepMind, as an IC via Topcoder, on RSI, Gemini evaluations, and Colab Bench and ML RL environments. Before that I was at Abundant AI (YC F24), building GPU-enabled RL environments for long-running agents, SWE and data-science environments, and the sandbox architecture underneath them. Some adversarial ML benchmarks I wrote there ended up in the training pipelines of three of the world's top five AI labs, which I still find slightly absurd.
The newest one, under review at ISEAI, is about reading multi-token answers out of a single hidden state: token lenses give you one token at a time and cannot spell "North Korea", so instead I patch a late-layer state into an early layer of a copy prompt and let the model say the answer itself. PEBS, at an ICML 2026 workshop, calibrates a reward model to each individual rater instead of to their average. Retroactive Advantage Correction, at the ICML 2026 RLxF workshop, is a closed-form fix for the case where the reward arrives after the step it belongs to, so you can train on it rather than drop it. A third paper, at the ICLR 2026 GRaM workshop, probes reasoning models with hyperbolic geometry and finds hierarchy in the deepest layers, where Euclidean probes lose it. I'm also a coauthor on KG-MuLQA, a long-context multi-hop QA benchmark at ACL 2026.
Earlier I was a research intern at Georgia Tech's Financial Services Innovation Lab, and at Harvard's Edge Computing Lab. I co-founded the AI Safety Club at IIT Delhi, and went through BlueDot Impact's alignment course and the ARENA curriculum.
Indian Institute of Technology Delhi
NK Security Scholar
Merit-based scholarship awarded to top 30 students at IIT Delhi for academic and technical excellence
Smart India Hackathon
National Top 5 Finalist in both 2023 and 2024 editions (India's largest student innovation competition)
JEE Advanced 2022
All India Rank 1,158 out of 1,000,000+ candidates (top 0.1%)
KVPY SX Fellowship 2021
National science fellowship awarded by the Government of India and IISc Bangalore
Codeforces Expert
1700+ rating in competitive programming
National Science Olympiads
Top 250 Astronomy, Top 300 Chemistry in India
Founding Member & Technical Lead / AI Safety Club, IIT Delhi
Co-founded student research group on AI alignment, interpretability, and evaluation. Led reading groups on mechanistic interpretability (TransformerLens) and safety frameworks. Completed BlueDot Impact AI safety training and ARENA curriculum.
Technical Consultant / STEM AI Hackathon 2026 (AI-Collab Hack)
Providing technical mentorship to 20+ teams building AI agents for STEM education at a hackathon jointly organized by IIT Delhi, Imperial College London, and Microsoft Garage.
Senior Editor / Tech Ambit (Pan-IIT Magazine)
Led 15-member editorial team across 23 IITs; curated and edited 30+ technical articles on AI and systems research.
Mess Secretary / Zanskar Hostel, IIT Delhi
Elected by 400+ residents; managed operations team of 13. Awarded Best Mess Secretary for digitalization initiatives.
My research, mapped: each colour is a domain (safety & eval, RL & training, foundations & language); the bridges show where those domains meet. Click any topic for the papers behind it.