publications

publications in reverse chronological order.

2026

  1. Ave: Guiding Agentic GPU Optimization using Data-Flow Invariants
    Haohui Mai, Xiaoyan Guo, Xiangyun Ding, and 7 more authors
    The 32nd ACM Symposium on Operating Systems Principles (SOSP 2026), to appear, Jul 2026
  2. When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output
    Shuoming Zhang, Jiacheng Zhao, Hanyuan Dong, and 9 more authors
    The ACM Conference on Computer and Communications Security (CCS 2026), to appear, May 2026
  3. Symbiotic MLLM Serving: Dynamically Balancing Parallelism Across GPUs and Resources Within GPUs
    Zhicheng Li, Jiacheng Zhao, Yangyu Zhang, and 10 more authors
    The 53rd Annual International Symposium on Computer Architecture (ISCA 2026), to appear, May 2026
  4. LEGO: An LLM-Enabled Hierarchical Optimizer for Tensor Computation Graphs with Structure-Aware Search and Compositional Synthesis
    Ruiyuan Xu, Shuoming Zhang, Guangli Li, and 10 more authors
    International Conference on Machine Learning (ICML 2026), to appear, May 2026
  5. CONTINUUM: Restoring the Contiguous Tensor Abstraction Efficiently for Dynamic AI Workloads via Hardware Virtualization
    Yangyu Zhang, Shuoming Zhang, Chunwei Xia, and 10 more authors
    International Conference on Machine Learning (ICML 2026), Spotlight (top 2.2%), to appear, May 2026
  6. A GPU Memory Allocator with Device-Side Page Table Materialization and Deferred TLB Coherence
    Yangyu Zhang, Lei Chen, Chunwei Xia, and 12 more authors
    USENIX Symposium on Operating Systems Design and Implementation (OSDI 2026), to appear, May 2026
  7. From Threads to Tiles: T2T, a Compiler for CUDA-to-NPU Translation via 2D Vectorization
    Shuaijiang Li, Jiacheng Zhao, Ying Liu, and 11 more authors
    The 24th ACM/IEEE International Symposium on Code Generation and Optimization (CGO 2026), Distinguished Paper Award, Feb 2026

2025

  1. SpaceServe: Spatial Multiplexing of Complementary Encoders and Decoders for Multimodal LLMs
    Zhicheng Li, Shuoming Zhang, Jiacheng Zhao, and 8 more authors
    The Thirty-ninth Annual Conference on Neural Information Processing Systems (NeurIPS 2025), Dec 2025
  2. Beyond Prompts: Space-Time Decoupling Control-Plane Jailbreaks in LLM Structured Output
    Shuoming Zhang, Jiacheng Zhao, Hanyuan Dong, and 5 more authors
    arXiv preprint arXiv:2503.24191 (2025), Mar 2025
  3. TopServe: Task-Operator Co-scheduling for Efficient Multi-DNN Inference Serving on GPUs
    Ao Chen, Guangli Li, Feng Yu, and 5 more authors
    European Conference on Parallel Processing, pp. 292-305. Cham: Springer Nature Switzerland, 2025, Jan 2025
  4. Qiwu: Exploiting Ciphertext-Level SIMD Parallelism in Homomorphic Encryption Programs
    Zhongcheng Zhang, Ying Liu, Yuyang Zhang, and 5 more authors
    Proceedings of the 23rd ACM/IEEE International Symposium on Code Generation and Optimization (CGO 2025), pp. 523-537, Jan 2025
  5. Fast and scalable neural network quantum states method for molecular potential energy surfaces
    Yangjun Wu, Wanlu Cao, Jiacheng Zhao, and 1 more author
    IEEE Transactions on Parallel and Distributed Systems (2025), Jan 2025

2024

  1. Introducing compiler semantics into large language models as programming language translators: A case study of C to x86 assembly
    Shuoming Zhang, Jiacheng Zhao, Chunwei Xia, and 3 more authors
    Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 996-1011, Nov 2024
  2. Optimizing Dynamic-Shape Neural Networks on Accelerators via On-the-Fly Micro-Kernel Polymerization
    Feng Yu, Guangli Li, Jiacheng Zhao, and 3 more authors
    29th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS 2024), Apr 2024
  3. Optimizing Deep Learning Inference via Global Analysis and Tensor Expressions
    Chunwei Xia, Jiacheng Zhao, Qianqi Sun, and 5 more authors
    29th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS 2024), Apr 2024
  4. Enabling Large Dynamic Neural Network Training with Learning-based Memory Management
    Jie Ren, Dong Xu, Shuangyan Yang, and 6 more authors
    30th IEEE International Symposium on High-Performance Computer Architecture (HPCA 2024), Mar 2024

2023

  1. Honeycomb: Secure and Efficient GPU Executions via Static Validation
    Haohui Mai, Jiacheng Zhao, Hongren Zheng, and 7 more authors
    7th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23), Jul 2023
  2. Sirius: Harvesting Whole-Program Optimization Opportunities for DNNs
    Yijin Li, Jiacheng Zhao, Qianqi Sun, and 12 more authors
    Sixth Conference on Machine Learning and Systems (MLSYS), Jun 2023
  3. PosFuzz: Augmenting Greybox Fuzzing with Effective Position Distribution
    Yanyan Zou, Wei Huo, Jiacheng Zhao, and 3 more authors
    Cybersecurity, Jan 2023

2022

  1. Unified holistic memory management supporting multiple big data processing frameworks over hybrid memories
    Lei Chen, Jiacheng Zhao, Chenxi Wang, and 8 more authors
    ACM Transactions on Computer Systems, Jul 2022
  2. VTensor: Using Virtual Tensors to Build a Layout-Oblivious AI Programming Framework (Full Version)
    Feng Yu, Jiacheng Zhao, Huimin Cui, and 2 more authors
    Journal of Computer Science and Technology, Apr 2022

2020

  1. VTensor: Using Virtual Tensors to Build a Layout-Oblivious AI Programming Framework (Poster Version)
    Feng Yu, Jiacheng Zhao, Huimin Cui, and 2 more authors
    PACT ’20: Proceedings of the ACM International Conference on Parallel Architectures and Compilation Techniques, Sep 2020

2019

  1. DNNTune: Automatic benchmarking DNN models for mobile-cloud computing
    Chunwei Xia, Jiacheng Zhao, Huimin Cui, and 2 more authors
    ACM Transactions on Architecture and Code Optimization, Dec 2019

2018

  1. On retargeting the ai programming framework to new hardwares
    Jiacheng Zhao, Yisong Chang, Denghui Li, and 4 more authors
    Network and Parallel Computing: 15th IFIP WG 10.3 International Conference, NPC 2018, Nov 2018
  2. Characterizing DNN models for edge-cloud computing (Poster)
    Chunwei Xia, Jiacheng Zhao, Huimin Cui, and 1 more author
    2018 IEEE International Symposium on Workload Characterization (IISWC), Sep 2018
  3. Revisiting loop tiling for datacenters: live and let live
    Jiacheng Zhao, Huimin Cui, Yalin Zhang, and 2 more authors
    ICS ’18: Proceedings of the 2018 International Conference on Supercomputing, Jun 2018

2016

  1. Predicting cross-core performance interference on multicore processors with regression analysis
    Jiacheng Zhao, Huimin Cui, Jingling Xue, and 1 more author
    IEEE Transactions on Parallel and Distributed Systems , May 2016

2015

  1. Hadoop+: Modeling and Evaluating the Heterogeneity for MapReduce Applications in Heterogeneous Clusters
    Wenting He, Huimin Cui, Binbin Lu, and 7 more authors
    Proceedings of the 29th ACM on International Conference on Supercomputing (ICS), Jun 2015

2013

  1. An empirical model for predicting cross-core performance interference on multicore processors
    Jiacheng Zhao, Xiaobing Feng, Huimin Cui, and 3 more authors
    Proceedings of the 22nd international conference on Parallel architectures and compilation techniques (PACT), Sep 2013