Hi, Iβm Runguo Li π
I am an M.S. student in Information Science at the University of Illinois Urbana-Champaign (UIUC).
I work on ML systems for large-scale models: inference engines and expert streaming for MoE models, distributed training and rollout infrastructure, RL post-training systems, and GPU compiler and kernel-level correctness. I like to sit between the machine and the model β how weights are read from storage, how collectives and thread teams are sized, and whether a resource plan and the runtime agree about what will actually be built. I keep a research background in LLM reasoning, multimodal learning, and agent safety, and I bring an upstream-first habit to all of it: reproduce failures in real production stacks, then contribute fixes with regression tests and measured before/after numbers.
My research experience includes research at UIUC on efficient ML systems, the Tencent Youtu AI Lab on content-safety and multimodal research, the SUFE FinAI Center (advised by Prof. Liwen Zhang) on financial reasoning and agent safety, Shanghai Jiao Tong University (advised by Prof. Xiaodong Gu) on LLM for Code, and the Head Office of ICBC (Private Banking Department) on scientist-discovery agents.
Research Interests
- ML Systems for Large Models β inference engines and expert streaming for MoE models, KV-cache and memory budgeting, attention and quantization kernels, performance profiling
- Distributed Training & Rollout β ZeRO-3, FSDP, Megatron-LM, Hybrid Engine, collective synchronization and deadlock debugging, RL rollout systems (verl, TRL)
- GPU Compiler & Kernel Correctness β MLIR passes, Triton/CUDA kernel-level debugging, reproducible regression tests
- LLM Reasoning & Multimodal Learning β chain-of-thought supervision, selective fine-tuning, retrieval and fusion
- Agents & Security β agentic planning and tool use, RAG and long-term memory, execution-grounded evaluation, runtime safety gates
Selected Highlights
- π FinVault β Benchmarking Financial Agent Safety in Execution-Grounded Environments β arXiv preprint, co-first author. arXiv:2601.07853
- π RVCFT β Reasoning-Visual Critical Token Fine-Tuning for Multimodal Reasoning β co-first author, under review at AAAI 2027.
- π VeriBRT β Plan-Guided, Evidence-Based Automated Bug Reproduction β co-first author, under review at ICSE 2027.
- π οΈ Open-source ML systems β merged fixes in PyTorch, DeepSpeed, colibri, FlashAttention, verl, Triton, Apache TVM, LMCache, and vLLM-Omni, plus bug reports fixed upstream in MLflow. See the full portfolio
- π οΈ ARH (AI Research Helper) β open-source CLI-first research assistant agent with tool use, plan/execute safety, three-layer memory, skill self-learning and multi-LLM fallback. github.com/LiRunGuo/Arhelper
Open-Source Contributions (ML Systems)
Merged fixes to PyTorch, DeepSpeed, a pure-C MoE inference engine, FlashAttention, verl, Triton, Apache TVM, LMCache, and vLLM-Omni, plus an MLflow bug report whose fix landed upstream. The list is ordered by the upstream projectβs star count.
- PyTorch β CUDA kernels β fixed an int32 overflow in the
cdistbackward kernel, where indices wrapped negative past 2Β³ΒΉ and slipped past the bounds check into an illegal memory access. - DeepSpeed β distributed training β fixed a ZeRO-3 rollout deadlock caused by unsynchronized generation stopping, and blocked partial Hybrid Engine policy injection for unsupported architectures (validated on H200 and MI250 GPUs).
- colibri β pure-C MoE inference engine β multi-drive expert streaming for the 510 GB DeepSeek-V4.1 container, physical-core OpenMP team sizing in four engines that were 18.7Γ slower without it, and a plan/runtime mismatch that silently allocated 6.25 GiB of unbudgeted KV cache.
- MLflow β SQLAlchemy stores β root-caused the numeric-attribute search failures under PostgreSQL with psycopg v3 to a string return in
SearchUtils, and the fix landed upstream (#26180). - FlashAttention β CuTe kernels β fitted the SM90 forward tile to the block-sparse block size (block sparsity on Hopper went from 8 of 40 head-dim/block-size combinations working to 32), and stopped the causal forward from re-applying an all-true mask on unmasked KV blocks, cutting its instruction count from 580 to 477 per loop.
- verl β RL post-training β restored FSDP value-head critic loading after TRL relocated its value-head model classes.
- Triton β GPU compiler β stopped nested-loop fusion from trusting an
llvm.assumeoutside the loop, which removed a zero-trip guard and could cause incorrect memory writes (MLIR regression test and H200 reproducer). - Apache TVM β ONNX frontend β let Hardmax lower when the reduced axis has a symbolic extent, instead of failing on
one_hotβs static-onlydepthattribute. - LMCache β KV cache management β made the allocator reject invalid sizes instead of silently emptying the cache, and enforced lazy
%-format logging (ruff G004) so new f-string logging can no longer land in already-migrated files. - vLLM-Omni β omni-modal inference β turned a malformed Qwen2.5-Omni prompt from an engine-killing crash into a rejected request.
Browse the full portfolio β fourteen merged pull requests across nine upstream projects, each with a reproducer or regression test attached.
Technical Skills
- ML Systems and Inference β MoE expert streaming and disk-resident inference, KV-cache and memory budgeting, attention and quantization kernels, serving engines (vLLM, SGLang, TensorRT-LLM), CUDA/Triton kernel-level debugging, GPU compiler passes (MLIR), performance profiling
- Training and Post-Training Infrastructure β distributed data/tensor/pipeline parallelism, ZeRO-3 and FSDP, Megatron-LM, DeepSpeed Hybrid Engine, collective synchronization and deadlock debugging, RL rollout systems (verl, TRL), SFT and preference optimization (DPO/GRPO/PPO), distillation, LoRA/PEFT
- Languages and Tooling β Python, C, CUDA/Triton, SQL, Bash, LaTeX; PyTorch, Transformers, FlashAttention, FAISS; Git-based upstream contribution, regression testing, reproducible benchmarking, Linux, Docker
- Systems for Agents β agentic planning and tool use, RAG and long-term memory, execution-grounded evaluation, runtime safety gates, multi-channel gateways (FastAPI)
News
- 2026.09 β Open-source work merged across PyTorch, DeepSpeed, colibri, FlashAttention, verl, Triton, Apache TVM, LMCache, and vLLM-Omni, plus a bug report fixed upstream in MLflow; portfolio now lists fourteen merged pull requests and is ordered by upstream star count.
- 2026.08 β Started the M.S. in Information Science program at the University of Illinois Urbana-Champaign.
- 2026.07 β Completed research internships at the SUFE FinAI Center and Shanghai Jiao Tong University.
- 2026.05 β Launched this personal homepage at runguoli.com. π
- 2026.03 β Joined Shanghai Jiao Tong University as a research intern (LLM for Code, advised by Prof. Xiaodong Gu).
- 2026.01 β FinVault released as an arXiv preprint; started working with the Head Office of ICBC on scientist-discovery agents.
- 2025.10 β Joined Tencent Youtu AI Lab as a research intern.
Get in Touch
- βοΈ Email: runguo.ai@gmail.com
- π GitHub: github.com/LiRunGuo
- π Curriculum Vitae
