Rohit Yelukati Mahendra PRO
AI & ML interests
Recent Activity
Organizations
Posts 1
I fine-tuned OpenBMB's MiniCPM5-1B to write Triton GPU kernels, then let an immutable referee decide if they are real: compile, check correctness against PyTorch on adversarial inputs, time against eager, torch.compile, and torch.compile max-autotune, then block the known ways of gaming the benchmark.
The 1B setup beat torch.compile max-autotune in 12/12 independently seeded runs. The larger Qwen3.6-27B smith pushed the same referee loop further: 76 verified compiler-beating kernels on H200, with 69 surviving a 5-run stability gate and 7 kept as single-shot probes on unseen problems. On a 376-cell shape/dtype grid, the stability-gated kernels keep a 1.49x geomean, with about 10% of cells losing and reported per cell.
Honest bound: these are scheduling wins on memory-bound ops, not new algorithms or wins over cuBLAS/FlashAttention. The scarce thing is not the big model, it is the verifier it cannot fool.
Full write-up: https://huggingface.co/blog/YMRohit/ouroboros-kernel-mint
Try it: build-small-hackathon/ouroboros-kernel-mint
2-min demo: https://youtu.be/ViicZHktb-A
Built for #BuildSmallHackathon with MiniCPM, Qwen, Triton, Gradio, Codex, and Modal H200s.
- Running
Repro - d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation
🎯Browse and share experiment logs with a coding agent
-
YMRohit/icml2026-61269-d3llm-repro-results
Updated • 100 - Running
Repro - MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
🎯Explore benchmark logs and sync findings with an AI agent
-
YMRohit/icml2026-64918-memorybench-repro-results
Preview • Updated • 125
- Running
Repro - Optimal Unconstrained Self-Distillation in Ridge Regression
🔬Explore and collaborate on project logbooks online
-
YMRohit/icml2026-22249-self-distillation-logbook-artifacts
11.7 MB - PausedAgents
Icml22249 Release Artifacts
🎯Display a visual summary of your program's I/O activity
- Running
Repro - d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation
🎯Browse and share experiment logs with a coding agent
-
YMRohit/icml2026-61269-d3llm-repro-results
Updated • 100 - Running
Repro - MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
🎯Explore benchmark logs and sync findings with an AI agent
-
YMRohit/icml2026-64918-memorybench-repro-results
Preview • Updated • 125
- Running
Repro - Optimal Unconstrained Self-Distillation in Ridge Regression
🔬Explore and collaborate on project logbooks online
-
YMRohit/icml2026-22249-self-distillation-logbook-artifacts
11.7 MB - PausedAgents
Icml22249 Release Artifacts
🎯Display a visual summary of your program's I/O activity
spaces 26
Repro - On Uniform Error Bounds for Kernel Regression under Non-Gaussian Noise
Explore and edit a collaborative research logbook
Icml27364 Kernelbounds
Display experiment tracking results in an interactive view
Repro - Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
Organize research notes and sync with an AI coding agent
Icml5886 Softmaxlinear
Show a visual timeline of your program's I/O activity
Repro - Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad
Browse experiment logs and sync findings with an AI agent
Icml31042 Adagradheavy
Show an interactive tracking dashboard for your data