publications
publications in reverse chronological order.
2026
-
Ave: Guiding Agentic GPU Optimization using Data-Flow InvariantsThe 32nd ACM Symposium on Operating Systems Principles (SOSP 2026), to appear, Jul 2026
-
When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured OutputThe ACM Conference on Computer and Communications Security (CCS 2026), to appear, May 2026
-
Symbiotic MLLM Serving: Dynamically Balancing Parallelism Across GPUs and Resources Within GPUsThe 53rd Annual International Symposium on Computer Architecture (ISCA 2026), to appear, May 2026
-
LEGO: An LLM-Enabled Hierarchical Optimizer for Tensor Computation Graphs with Structure-Aware Search and Compositional SynthesisInternational Conference on Machine Learning (ICML 2026), to appear, May 2026
-
CONTINUUM: Restoring the Contiguous Tensor Abstraction Efficiently for Dynamic AI Workloads via Hardware VirtualizationInternational Conference on Machine Learning (ICML 2026), Spotlight (top 2.2%), to appear, May 2026
-
A GPU Memory Allocator with Device-Side Page Table Materialization and Deferred TLB CoherenceUSENIX Symposium on Operating Systems Design and Implementation (OSDI 2026), to appear, May 2026
-
From Threads to Tiles: T2T, a Compiler for CUDA-to-NPU Translation via 2D VectorizationThe 24th ACM/IEEE International Symposium on Code Generation and Optimization (CGO 2026), Distinguished Paper Award, Feb 2026
2025
-
SpaceServe: Spatial Multiplexing of Complementary Encoders and Decoders for Multimodal LLMsThe Thirty-ninth Annual Conference on Neural Information Processing Systems (NeurIPS 2025), Dec 2025
-
Beyond Prompts: Space-Time Decoupling Control-Plane Jailbreaks in LLM Structured OutputarXiv preprint arXiv:2503.24191 (2025), Mar 2025
-
TopServe: Task-Operator Co-scheduling for Efficient Multi-DNN Inference Serving on GPUsEuropean Conference on Parallel Processing, pp. 292-305. Cham: Springer Nature Switzerland, 2025, Jan 2025
-
Qiwu: Exploiting Ciphertext-Level SIMD Parallelism in Homomorphic Encryption ProgramsProceedings of the 23rd ACM/IEEE International Symposium on Code Generation and Optimization (CGO 2025), pp. 523-537, Jan 2025
-
Fast and scalable neural network quantum states method for molecular potential energy surfacesIEEE Transactions on Parallel and Distributed Systems (2025), Jan 2025
2024
-
Introducing compiler semantics into large language models as programming language translators: A case study of C to x86 assemblyFindings of the Association for Computational Linguistics: EMNLP 2024, pp. 996-1011, Nov 2024
-
Optimizing Dynamic-Shape Neural Networks on Accelerators via On-the-Fly Micro-Kernel Polymerization29th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS 2024), Apr 2024
-
Optimizing Deep Learning Inference via Global Analysis and Tensor Expressions29th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS 2024), Apr 2024
-
Enabling Large Dynamic Neural Network Training with Learning-based Memory Management30th IEEE International Symposium on High-Performance Computer Architecture (HPCA 2024), Mar 2024
2023
-
PosFuzz: Augmenting Greybox Fuzzing with Effective Position DistributionCybersecurity, Jan 2023