Search

Search articles, pages, topics, and people.

News

Recent publications, awards, and other updates.

2026

  1. Publication: Our paper “GauTracer” is accepted by ISCA 2026. Looking forward to sharing our work at Raleigh, North Carolina!
    GauTracer fully integrates Gaussian primitives into the GPU Ray Tracer pipeline, eliminating software shader invocations.
  2. Publication: Our paper “SpikeLet” was accepted by TCAD.
    This paper presents SpikeLet, a event-driven spatial-temporal-parallel multichiplet neuromorphic system tailored for large-scale SNNs.
  3. Publication: Our paper “CIM-Pruner” is accepted by ISCAS 2026.
    This paper demonstrates a Compute-in-Memory macro design supporting in-memory token merging and pruning for Vision-Language Models.

2025

  1. Award: Our paper won IEEE A-SSCC 2025 Distinguished Design Award!
    It is our second time to win the Distinguished Design Award on IEEE A-SSCC. And two students, Siqi He and Yujie Ma, received the IEEE WiC TG Award.
  2. Award: Our paper “GauRPE” is selected as ICCAD 2025 Best Paper Award Candidate!
    GauRPE is a SW/HW co-design that performs 3DGS rasterization via pattern matching.
  3. Publication: Our ISSCC'25 paper about “SHINSAI: Reusable Active TSV-Interposer” was invited to JSSC and is accepted.
    SHINSAI (芯斋) is a 586mm2 reusable active TSV interposer with microbump-level programmable interconnect fabric and 512Mb 3D underdeck SRAM memory.
  4. Publication: Our paper about “VLA-driven Manipulator on FPGA” is accepted by A-SSCC 2025 and selected as a Highlight paper.
    We will demonstrate the robotic arm driven by our FPGA accelerator at the Demonstration Session. We look forward to seeing you in Korea.
  5. Publication: Our paper about “Pattern-based Rendering Engine for 3DGS” is accepted by ICCAD 2025.
    GauRPE is a SW/HW co-design that performs 3DGS rasterization via pattern matching. Looking forward to the presentation in Munich.
  6. People: I became an tenure-track Assistant Professor at Fudan University.
    Still focusing on AI chips for Robotics/LLM/.... Welcome to apply for doctoral and master's degrees!
  7. Publication: Our paper “A 22-nm 109.3-to-249.5-TFLOPS/W Outlier-Aware ...” is accepted by JSSC.
    This paper proposes OA-CIM, a 22nm SRAM-based Compute-in-Memory macro for LLMs. It supports mixed-precision (INT4+FP16) MACs optimized for outlier-aware LLM deployment.
  8. Publication: Our 2 papers about “HW/SW-codesign for MoE” are accepted by DAC 2025.
    PIMoE proposes a workload offloading strategy for MoE deployment on NPU-PIM heterogeneous systems, while Hydra addresses the load imbalance challenge among MoE experts in multi-chiplet integrated chips.
  9. Publication: Our paper “SHINSAI: A 586mm2 Reusable Active TSV-Interposer ...” appeared on ISSCC 2025.
    SHINSAI (芯斋) is a 586mm2 reusable active TSV interposer with microbump-level programmable interconnect fabric and 512Mb 3D underdeck SRAM memory.

2024

  1. Award: Our paper “A Real-Time Optical-Flow-based SLAM ...” won the Distinguished Design Award on A-SSCC 2024!
    We presented and demonstrated Xiliu, an optical-flow-based SLAM accelerator on FPGA. It exploits the similarity and sparsity of optical flow to achieve real-time performance.
  2. Publication: Our paper “GauSPU: 3D Gaussian Splatting Processor ...” appeared on MICRO 2024.
    GauSPU is a HW/SW-cooptimized GPU extension aiming to realize real-time pose tracking in 3D Gaussian Splatting-based SLAM systems.
  3. Publication: Our paper “ST-BPTT: a Memory-Efficient BPTT SNN Training ...” appeared on BioCAS 2024.
    This work reduces the memory footprint of BPTT-based SNN training by cutting off the back-propagation paths of the timesteps with low significance.
  4. Publication: Our 2 Compute-in-Memory papers appeared on ISCAS 2024.
    We presented two works, a Logarithmic FP CIM macro and a LUT-based CIM macro. The latter, Trident-CIM, was invited to transfer to TCAS-II.
  5. Publication: Our paper “SLAM-CIM: A Visual SLAM Backend Processor ...” is accepted by JSSC.
    We propose SLAM-CIM, a Compute-in-Memory based visual SLAM backend processor for edge robotics. It features SRAM-based FP16 digital CIM macros supporting both MAC and linear solving.