Electrical Engineering · IIT Bombay

Arnav Agarwal

I build at the line where machine learning meets silicon, where a model stops being a notebook experiment and has to run inside a power budget.

Arnav Agarwal
Program
B.Tech + M.Tech, EE · 2022–2027
Focus
ML · Digital systems · Edge AI
Publications
2 × IEEE · AICAS, APCCAS 2026
Based in
Mumbai, India

What I’m after

I like problems that don’t let you cheat. A network that has to fit in thirty kilobytes, a pipeline that has to keep pace with a live camera, a chip that either closes timing or doesn’t. The constraint is the interesting part, because it forces you to understand the thing rather than throw capacity at it.

What keeps me in a project is seeing it all the way through. A result I can measure is good; a working thing someone can pick up and use is better. That’s why I keep ending up on problems where the research and the engineering aren’t really separable, where the idea only counts once the hardware runs. I’d rather spend a year making one thing genuinely work than move quickly across things that never leave the bench.

So what I’m looking for next is an environment with real technical depth and a short path from an idea to something real: a group where the work is meant to end up in the world, in whatever form that takes. I’m still figuring out the specifics, which is a fair part of why this page exists.

If you’ve taken deep technical work and turned it into something people actually use, whether in a lab, a company, or somewhere in between, I’d genuinely like to hear how that went. And if you’re working anywhere near this intersection, I’d like to hear about that too.

Now

Thesis

Designing a chip that will go to fabrication: a vision processor built from the logic up, working toward the version that gets manufactured.

Under review

Submitted a paper on inferring stochastic differential equations to the NeurIPS 2026 STODY workshop, from my work at IBM Research.

Presenting

Two papers at IEEE AICAS and APCCAS 2026, both as lecture presentations.

Travelling

I’ll be at APCCAS in Fukuoka, Japan, 25–28 October. If you’re going to be there too, I’d like to say hello.

Teaching

TA for EE340 Communications Lab at IIT Bombay, running weekly GNURadio SDR and hardware sessions.

Selected work

Silicon & systems

Processors, accelerators, and custom chips. Once something is manufactured you can’t patch it, which changes how you build.

Streaming Vision SoC

M.Tech thesis, IIT Bombay · 2025–present

Algorithm-and-silicon co-design for a wearable vision implant. A fixed-point streaming front-end that never buffers a full camera frame, so the entire seven-stage pipeline runs in 8.56 kB of on-chip memory. On top of it, NanoSeg: a segmentation CNN designed from scratch across seven documented iterations to run within the memory left over, 12.6 kB of INT8 weights and 28.8 kB total. Full RTL-to-GDS2 in TSMC 65 nm; two IEEE papers.

  • 65 nm TSMC
  • 0.81 mm²
  • 9.5 mW
  • 8.56 kB on-chip
  • 12.6 kB INT8 CNN
  • 21× ENet acc/param

Superscalar Out-of-Order Processor

EE739, IIT Bombay · 2026

A 2-way 16-bit superscalar core in VHDL with full out-of-order execution: reservation stations, reorder buffer, common data bus forwarding, and a two-level branch predictor.

  • 2 instr/cycle
  • 16-entry ROB
  • VHDL

Spiking Neural Prefetcher

CS683, IIT Bombay · 2025

Extended the PATHFINDER prefetcher in ChampSim with a spiking neural network gating predictions by confidence, falling back to load-value prediction when the SNN is unsure.

  • 2× SNN accuracy
  • 36K→14.5K preds

Real-Time ADAS Segmentation on FPGA

EE712, IIT Bombay · 2026

Road-scene segmentation deployed as custom programmable-logic IP on a PYNQ-Z2, with INT8 quantization cutting weight memory 4× to fit the on-chip budget.

  • 11K params
  • 87% BRAM
  • XC7Z020

Machine learning

Models and the systems that run them: generative modelling, inference performance, and problems where the data or the compute is the binding constraint.

Learning Stochastic Differential Equations

IBM Research India · 2026

A moment-field encoder with a region-based mixture-of-experts decoder that infers stochastic differential equations from observed trajectories, beating a transformer foundation-model baseline on every axis.

  • 27% lower RMSE
  • 4× fewer params
  • 4× faster

Generative 3D CAD

Arizona State University · 2025

Cascaded transformers with a pointer network mapping 2D images into editable B-rep CAD geometry, plus a construction engine that exports real, openable STEP files.

  • 0.72 BLEU-4
  • 94% STEP validity
  • 500K models

CUDA Kernels & LLM Inference

CS794, IIT Bombay · 2026

FlashAttention written from scratch in CUDA with online-softmax tiling, verified against PyTorch’s kernel, alongside KV caching and prefix sharing through a shared radix tree.

  • 6.8× matmul speedup
  • 3.5 TFLOPS

Generative Models for Few-Shot Medical Imaging

EE782, IIT Bombay · 2025

Trained CVAE, DCGAN and a U-Net diffusion model from scratch to synthesize dermoscopic images for rare skin-lesion classes, then measured what each actually bought in downstream accuracy.

  • +6.6 pts (diffusion)
  • 1,200 images

Publications

[1]

Hardware-Aware Semantic Segmentation Engine for On-Chip Retinal Prosthesis Navigation

IEEE AICAS 2026 · lecture presentation

Arnav Agarwal*, Daksh Sawke*, Laxmeesha Somappa  *equal contribution

Scene understanding designed to run entirely on the spectacle-mounted processor, with no phone and no cloud, inside the memory budget of a device worn on a pair of glasses. A 12.6 kB CNN that reaches 21× ENet’s accuracy per parameter.

[2]

An On-Chip Streaming Fixed-Point Vision Processing Pipeline for Epiretinal Prostheses

IEEE APCCAS 2026 · lecture presentation

Arnav Agarwal*, Daksh Sawke*, Laxmeesha Somappa  *equal contribution

A seven-stage pipeline that turns a live camera feed into retinal stimulation signals never buffering a full frame anywhere in the datapath. Implemented in TSMC 65 nm CMOS at 0.81 mm² and 9.5 mW, and evaluated on what the user would actually perceive rather than on the image alone.

Both papers are accepted and will be presented in 2026. The PDFs above are the accepted author versions.

Built & shipped

Work that made it out of the lab and in front of real users.

Generative AI Intern, Trupeer Technologies Private Limited

Built backend services and a Chrome extension that turns a rough screen recording into a finished product video and written user guide in under five minutes, cutting client onboarding time by 80%. Also built a multilingual document reader with text-to-speech across Indian languages, which was selected for presentation at UNESCO’s 46th World Heritage Committee Meeting.

shipped

Supply Chain Risk ML, GoComet India Private Limited

Built the ML system that reads 100,000+ news and RSS feeds a month and warns logistics customers which of their shipments a port disruption is about to hit, before it does. Ran in production on AWS at 85%+ precision and under two seconds.

in production

Buddy Booth, HKUST Entrepreneurship Bootcamp

An AI kiosk that picks up early signs of bullying in primary schools. Built and pitched it in a week in Hong Kong as one of three IIT Bombay nominees on India’s only delegation.

3rd of 15+

Searce × Google Cloud GenAI Hackathon

A bilingual app that diagnoses crop health from a photo a farmer takes, with weather-aware advice in Hindi or English. Built in 36 hours; we were the youngest team in a field of twenty industry teams.

2nd of 20

IDEAS incubation · DSSE, IIT Bombay

Selected for Cohort 13 of the institute’s startup incubation programme. Earlier, won ENPRO at E-Cell for a freight platform that cut empty return trips by 65%.

incubated

Experience

AI Researcher  IBM Research India

May–Aug 2026

Generative AI Research Intern  Arizona State University

May–Jul 2025

Machine Learning Engineer  GoComet India Private Limited

Feb–May 2025

Generative AI Engineer  Trupeer Technologies Private Limited

May–Jul 2024

Get in touch

Institute email22b3917@iitb.ac.in
RésuméDownload PDF
Based inMumbai, India