Program
Registration & Breakfast
Organizers: Zishen Wan, Chenyu Wang, Shvetank Prakash, Andy Cheng, Arya Tschand, Vijay Janapa Reddi (Harvard)
The Architecture 2.0 workshop aims to bring together researchers and practitioners exploring how AI can fundamentally transform the design, analysis, optimization, and evaluation of computer architectures and systems. We invite both work-in-progress and completed research that advances the state of the art or opens new research directions at the intersection of AI, architecture, and systems. Topics of interest include, but are not limited to:
- Computer Architecture
- AI-driven microarchitecture design, tuning, and exploration
- AI-driven chip design, hardware code generation, synthesis, and place-and-route
- Learning-based design space exploration and architectural trade-off analysis
- Processor, accelerator, and heterogeneous system design
- GPU, TPU, and accelerator architectures and kernel optimization
- Memory systems, cache hierarchies, interconnects, and storage optimization using AI
- Learning-based performance, power, energy, and reliability modeling
- Surrogate models to accelerate architectural simulation and evaluation
- AI methods for identifying architectural bottlenecks and inefficiencies
- Systems
- AI-assisted compilers, program analysis, and code generation
- Automatic kernel transformation, scheduling, and tuning
- Cross-layer co-design across compilers, runtimes, operating systems, and hardware
- AI-driven scheduling, resource management, and system-level optimization
- Intelligent runtime systems for heterogeneous and accelerator-rich platforms
- Learning-based memory management, caching, and I/O policies
We will also conduct a tutorial on QuArch and ArchEval – datasets and characterization of AI agents as computer architects. As LLM agents begin to propose architecture mechanisms, simulator configurations, accelerator mappings, and hardware–software co-design choices, the key question is no longer whether they can generate plausible artifacts, but whether they can reason like architects: analyze workloads, formulate design goals, use simulators, predict performance, satisfy hard constraints, and decide which feasible designs are worth evaluating. Using QuArch and ArchEval as a case study, the tutorial will present a benchmark methodology for measuring these capabilities across CPU cores, memory systems, ML accelerators, distributed systems, and compute-in-memory designs. We will discuss the L1/L2/L3 evaluation settings, which progressively remove prepared harnesses and simulator feedback to distinguish assisted design-space exploration from autonomous architectural judgment. The tutorial will also cover multi-simulator benchmarking, baseline-normalized scoring, trajectory analysis, and common failure modes of current agents.computer architectures and systems. We invite both work-in-progress and completed research that advances the state of the art or opens new research directions at the intersection of AI, architecture, and systems.
Organizers: Binuraj Ravindran, Rakib Al-Fahad (Intel)
Memory capacity is increasingly becoming a limiting factor for modern cloud infrastructure, in-memory databases, and AI-serving workloads, making memory compression an attractive solution for increasing effective memory capacity and improving resource utilization. This tutorial provides a practical introduction to compressed memory systems in Linux, focusing on zswap and zram, and explores how hardware accelerators such as Intel® In-Memory Analytics Accelerator (IAA) can reduce compression overhead while enabling higher workload density. Participants will learn the fundamentals of compressed memory, Linux memory-management internals, and methodologies for evaluating key metrics including compression ratio, memory savings, CPU utilization, swap latency, throughput, and tail latency. Through case studies using Redis and other memory-intensive workloads, the tutorial demonstrates how workload characteristics influence the effectiveness of memory compression and highlights the tradeoffs between capacity expansion and application performance. By combining operating-system internals, workload characterization techniques, and real-world deployment experiences, this tutorial equips researchers and practitioners with the knowledge needed to evaluate and deploy compressed memory systems in modern data-center environments.
Coffee Break
Organizers: Ben Feinberg and William Chapman (Sandia National Laboratories)
Understanding the impacts of analog nonidealities is one of the most important aspects of designing novel hardware based on analog crossbars. Ideas which can seem like straight-forward transformations can yield dramatically different susceptibility to analog errors. This tutorial will cover the use of CrossSim to simulate a range of analog nonidealities in neural networks, digital signal processing, and scientific computing applications. This tutorial will also cover how hardware models can be added to CrossSim for use beyond architecture research. In addition to covering CrossSim for conventional analog crossbar accelerators, this talk will also discuss how CrossSim can be used to design and model analog architectures based on spiking and neuromorphic concepts using the same core non-ideality simulation.
Lunch
Organizers: Dima Nikiforov, Shengjun Kris Dong, Agustin Coppari Hollmann, Loren Hung, Ailsa Sun, Yakun Sophia Shao (UC Berkeley)
This tutorial presents an open-source runtime and scheduling stack for building, tracing, and optimizing end-to-end ML and robotic workloads on heterogeneous RISC-V SoCs. The stack combines a lightweight Zephyr-RTOS ML runtime, an end to end PyTorch model compilation flow, and an orchestration layer for multi-rate real-time workloads.
First, participants will learn how to compile PyTorch models ahead-of-time into compact Zephyr ELFs with target-specific kernels, static tensor storage, persistent hart-pinned worker pools, explicit core affinity, accelerator dispatch, correctness validation, and first-class per-dispatch tracing. The runtime supports workloads running across scalar RISC-V cores, vector processors, tightly coupled memories, and custom accelerators such as systolic arrays and outer-product units.
Secondly, the tutorial addresses the full-system problem of deploying a multi-model workload under real-time constraints to heterogeneous SoCs. This section will build upon a robotics workload with perception, planning, control, and vision-language-action (VLA)-style components. We cover how to use backend choice, task fidelity, granularity, and execution frequency as control knobs, and how to assign temporal constraints and optimization targets. Additionally, we walk through how to use an ahead-of-time scheduler to build schedules that bind model operations to HW backends, and deploy these schedules to the RTOS-based runtime.
Finally, Participants will take a VLA robotic workload from PyTorch to a running, traced, schedule-driven binary that can be deployed across instruction-set simulation, software RTL simulation, FPGA-accelerated simulation with FireSim, and commercial heterogeneous silicon. The same runtime flow can be used to study accelerator integration, multicore scheduling, memory hierarchy design, hardware/software co-design, and performance analysis pre-silicon.
The tutorial also shows how runtime traces can feed autotuning and agentic optimization loops that generate, validate, and iteratively optimize kernels and placement decisions based on measured hardware behavior. It concludes with a discussion of future lightweight ML runtimes for heterogeneous embedded systems, emphasizing ahead-of-time model compilation, explicit hardware topology, portable accelerator interfaces, first-class tracing, and closed-loop optimization.
Organizers: Mahesh Madhav, Kunal Kashyap (two current members of SPEC CPU committee)
In May 2026, the latest version of SPEC CPU was released (CPU 2026). As leaders of the committee that crafted this benchmark suite, we want to provide a “behind the scenes” look at the process. We will share how applications were adapted into candidate benchmarks, and how the candidates were culled to create the final benchmark suite. We will share details on why certain classes of applications were excluded, and what that means. We will include a tutorial on recurrence plots for basic-block-vectors, which are a visual aid for determining code self-similarity within a benchmark, one of the inputs for benchmark selection. We will also touch upon the new heterogeneous schedule called rolling round-robin (RRR). There are more multithreaded benchmarks than ever before, so we will share early learnings from MT analysis of the Speed benchmarks, plus other system level sensitivity analysis taken from many-core machines. We will discuss the adaptation fidelity of the SPEC CPU benchmarks, that is, how close do the benchmarks represent their original applications. We will end with how the community can continue to support and influence SPEC CPU.
1:30p - 1:40p Introduction
1:40p - 3:00p “SPEC CPU: Behind the Scenes” (Mahesh Madhav)
3:00p - 3:30p “Adaptation Fidelity” (Doa’a Al-Otoom)
3:30p - 4:00p Coffee break
4:00p - 5:00p “Cornucopia of Characterization” (Kunal Kashyap)
5:00p - 5:15p Wrap-up / discussion
5:15p - 6:00p Buffer
Coffee Break
Registration & Breakfast
BranchSherpa: A Fast and Accurate Analytical Modeling Framework for Branch Predictors
Dongin Lee (National University of Singapore), Trevor E. Carlson (National University of Singapore, Google)
BitstreamZoo: A Benchmark Suite for Bitstream Workloads on Commodity CPUs and GPUs
Hongyuan Liu (Stevens Institute of Technology), Tianao Ge (The Hong Kong University of Science and Technology (Guangzhou)), Ming Li (Stevens Institute of Technology), James Contini (Stevens Institute of Technology)
Integration, Enhancements, and Evaluation of Memory Simulators
Pouya Esmaili-Dokht (Barcelona supercomputing center, Unversitat Politecnica de Catalunya), Arash Yadegari (Barcelona Supercomputing Center, Sharif University of Technology), Victor Xirau (Barcelona Supercomputing Center), Julian Pavon Rivera (Barcelona Supercomputing Center), Hamid Sarbazi-Azad (Sharif University of Technology, IPM), Adrian Cristal (Barcelona Supercomputing Center, Unversitat Politecnica de Catalunya), Eduard Ayguade (Universitat Politecnica de Catalunya, Barcelona Supercomputing Center), Petar Radojković (Barcelona Supercomputing Center)
PerfCount: A Performance Counter Dataset for CPU Prediction Problems
Matthew Barondeau (The University of Texas at Austin), Erika S. Alcorta (The University of Texas at Austin), Maya Koppikar (The University of Texas at Austin), Andreas Gerstlauer (The University of Texas at Austin)
Break
Session chairs: Adwait Jog (Univ of Virginia), Asit Mishra (NVIDIA)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
Euijun Chung (Georgia Tech), Yuxiao Jia (Georgia Tech), Aaron Jezghani (Georgia Tech), Hyesoon Kim (Georgia Tech)
Compute-Communication Overlap Is Not Free: A Cross-Layer Characterization in GPU LLM Workloads
Jihwan Oh (Georgia Tech), Junkyum Kim (Georgia Tech), Seokjin Go (Georgia Tech), Jongse Park (KAIST), Divya Mahajan (Georgia Tech)
Characterizing How Complex Agentic AI Systems Handle General Tasks: A Trace-Based Simulation Study
Donghwan Kim (Penn State), Prakhar Singh (Penn State), Younghoon Min (SK Hynix), Jongryool Kim (SK Hynix), Jongse Park (KAIST), Kiwan Maeng (Penn State)
Performance Characterization of SPEC CPU®2026 on AMD EPYC™ 9755 Processor
Kunal Kashyap (AMD), Rajiv Ramanathan (AMD), Shayantika Bhattacharya (AMD)
HammerSim: A System-Level Tool to Model RowHammer
Kaustav Goswami (UC Davis), Ayaz Akram (UC Davis), Hari Venugopalan (UC Davis), Jason Lowe-Power (UC Davis)
Dissecting GPU Utilization for LLM Inference on Nvidia Hopper (not presented live)
Mohammad Siavashi (KTH Royal Institute of Technology), Gerald Q. Maguire Jr. (KTH Royal Institute of Technology), Dejan Kostic (KTH Royal Institute of Technology), Marco Chiesa (KTH Royal Institute of Technology)
To Persist or To Let Go? Characterizing The Workload-Dependent Trade-off of Preserving Progress in Intermittently-Powered NVFPGAs (not presented live)
Aalaa M.A. Babai (Kyushu University), Koji Inoue (Kyushu University)
Best Paper Award Ceremony (15 min)
Lunch
MAGMO: Multi-Attribute Graphs for Multiobjective Graph Processing with Application to Maritime Routing
Leo Gold (University of Connecticut), Afif Arif Siddiqi (University of Connecticut), David Sidoti (US Naval Research Laboratory), Krishna R Pattipati (University of Connecticut), Omer Khan (University of Connecticut)
ZRay: Portable Compiler-Assisted Memory Traffic Characterization
Ashwin Poduval (University of Wisconsin-Madison), Hayden Coffey (University of Wisconsin-Madison), Michael Swift (University of Wisconsin-Madison)
C2ME: Characterizing Ciphertext Matrix Multiplication Encodings for FHE Inference
Negar Neda (New York University), Akshath Mahajan (New York University), Brandon Reagen (New York University)
Characterizing Request and Token Energy Costs of LLM Inference Workloads on Modern CPU-GPU Systems
Prabhu Vellaisamy (Carnegie Mellon University), Vanessa Lam (Carnegie Mellon University), Shawn Blanton (Carnegie Mellon University), John Paul Shen (Carnegie Mellon University)
LifeKV: Reducing SSD Write Amplification with Lifetime-Aware KV Cache Placement in Multi-Turn LLM Serving
Junwoo You (Yonsei University), Junsung Kim (Yonsei University), Won Woo Ro (Yonsei University)
Activation Concentration: Characterizing Column-Level Output Sparsity Across Diffusion Model Architectures
Dazhi Yang (Fairleigh Dickinson University), Shafayat Mowla An (University of Colorado at Colorado Springs), Byeong Kil Lee (University of Colorado at Colorado Springs), Jeeho Ryoo (Fairleigh Dickinson University)
Break
FiCCO: Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
Shagnik Pal (University of Texas at Austin), Shaizeen Aga (AMD), Suchita Pati (AMD), Mahzabeen Islam (AMD), Lizy K. John (UT Austin)
Hardware Characterization of Diffusion vs. Autoregressive Language Model Inference: Compute vs. Memory Bottlenecks
Zhenxing Fan (University of Virginia), Kevin Skadron (University of Virginia)
MaestroRAG: Orchestrated Pipeline Architecture for Efficient RAG on Edge Devices
Deeksha Chaudhary (The Pennsylvania State University), Cyan Subhra Mishra (Arm Inc.), Rishabh Jain (Pennsylvania State University), Mahmut Kandemir (Penn State), Chita Das (Penn State University)
Anatomy of NEAR Sharded Blockchain: A Longitudinal Workload Characterization
Md. Ahsan Habib (Tulane University), Lu Peng (Tulane University)
Mind the Gap: The Disconnect Between Synthetic and Natural Edge Weights in Parallel Single-Source Shortest Path
Marco D’Antonio (Queen’s University Belfast), Thai Son Mai (Queen’s University Belfast), Hans Vandierendonck (Queen’s University Belfast)
TESSERA: A Workload-Driven Simulation and Design-Space Exploration Framework for Heterogeneous NPUs
Arghadip Das (Purdue University), Hoseok Kim (Purdue University), Soomin Lee (Purdue University), Arnab Raha (Intel Corporation), Deepak A Mathaikutty (Intel Corporation), Vijay Raghunathan (Purdue University)
Registration & Breakfast
Understanding and Exploiting Cache Asymmetry in Chiplet Processors via Profile-Predict-Guard Allocation
Sunwoo Kim (Ewha Womans University), Eunbi Jeong (Ewha Womans University), Myung Kuk Yoon (Ewha Womans University)
Understanding and Exploiting Diagonal Attention Sparsity in Autoregressive Image Generation
Daeun Kim (KAIST), Junwha Hong (Agency for Defense Development), Changhun Oh (KAIST), Yoonsung Kim (KAIST), Yoonhyeong Lee (Seoul National University), Jongse Park (KAIST)
VCSR-TC: A Sparse Matrix-Matrix Optimization Strategy for Leveraging Modern GPUs with Tensor Cores
Mouad Tiahi (Northeastern University), Elmira Karimi (Northeastern University), David Kaeli (Northeastern University)
When and Why to Use TMA: A Workload-Grounded Characterization of Tensor Memory Acceleration on NVIDIA Hopper GPUs
Christin David Bose (Purdue university), Cesar Avalos (Purdue University), Junrui Pan (Purdue University), Tim Rogers (Purdue University)
Agentic AI Workload Characterization
Yichao Yuan (University of Illinois Urbana-Champaign), Ankita Nayak (Gimlet Labs), Souvik Kundu (Intel Corporation), Nishil Talati (University of Illinois, Urbana Champaign)
A Benchmark for Cost and Energy-Efficient Execution of Neuroimaging Workflows on Commodity Clusters
Ajay Kumar (University of Missouri), Vladimir Omelyusik (University of Missouri), Satish Nair (University of Missouri), Praveen Rao (University of Missouri)
A Characterization of Google Workload Traces
Matthew Giordano (University of Washington), Akanksha Jain (Google), Enrico Deiana (Google), Derek Bruening (Google), Abhinav Sharma (Google), Ibrahim Hur (Google), Parthasarathy Ranganathan (Google), Thomas Anderson (University of Washington)
Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads
Yasmine Omri (Stanford University), Ziyu Gan (Independent), Zachary Broveak (Stanford University), Robin Geens (MICAS, KU Leuven), Zexue He (Stanford University), Alex Pentland (Stanford University, MIT), Marian Verhelst (MICAS, KU Leuven), Tsachy Weissman (Stanford University), Thierry Tambe (Stanford University)
Break
BloQBench: A Blockchain Benchmarking Framework for Quantum Supremacy
Nicholas Papadopoulos (University of Colorado Boulder), Ramin Ayanzadeh (University of Colorado Boulder)
CC-Bench: a Benchmark Suite for Collective Communication with Compression in HPC and AI Workloads
Wei Wang (National University of Defense Technology), Haozhe Fan (Nanjing University), Xingchen Liu (Institute of Computing Technology, Chinese Academy of Sciences), Man Liu (Institute of Computing Technology, Chinese Academy of Sciences), Xingjian Tian (Institute of Computing Technology, Chinese Academy of Sciences), Haoquan Long (Institute of Computing Technology, Chinese Academy of Sciences), Zedong Liu (Institute of Computing Technology, Chinese Academy of Sciences), Daran Sun (Institute of Computing Technology, Chinese Academy of Sciences), Jinwu Yang (Institute of Computing Technology, Chinese Academy of Sciences), Bo Yang (National University of Defense Technology), Jie Liu (National University of Defense Technology), Yonggang Che (National University of Defense Technology), Hairui Zhao (Institute of Computing Technology, Chinese Academy of Sciences), Guangming Tan (Institute of Computing Technology, Chinese Academy of Sciences), Dingwen Tao (Institute of Computing Technology, Chinese Academy of Sciences)
FHE-Lab: A Modular Framework for Customizable Fully Homomorphic Encryption Pipelines
Vattana Chan (Georgetown University), Matias Mazzanti (Universidad de Buenos Aires), Augusto Vega (IBM Research), Esteban Mocskos (Universidad de Buenos Aires), Radha Venkatagiri (Georgetown University)
From Chat to Agents on the Edge: A Cross-Framework Characterization of LLM Inference on Edge GPUs
Haebin Do (Harvard University), Kevin He (Harvard University), David Brooks (Harvard University), Gu-Yeon Wei (Harvard University)
Hydra: Phase-Aware Workload Characterization of LLM Inference across Edge SoC Generations, Backends, and Quantization Levels
Amir Taherin (Northeastern University), Sana Taghipour Anvari (Northeastern University), Charles Amante (Northeastern University), Yixiao Chen (Northeastern University), Ruben Noroian (Northeastern University), Zlatan Feric (Northeastern University), Nicolas Bohm Agostini (Northeastern University), Pu Zhao (Northeastern University), José Cano (University of Glasgow), Bin Ren (William & Mary), Yanzhi Wang (Northeastern University), David Kaeli (Northeastern University)
LLM-Quant-Lab: Characterizing Cross-Hardware Reproducibility of LLM Post-Training Quantization on AMD ROCm
Medhat Abouzeid (American University of Sharjah), Fadil Pasha (American University of Sharjah), Ihab Amer (American University of Sharjah)
PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving
Omkar Shewale (Illinois Tech), Deepak Kumar (Illinois Tech), Divakar Kumar Yadav (UWM Milwakee)
RAGMark: A Comprehensive Framework for Benchmarking Retrieval-Augmented Generation Systems
Zlatan Feric (Northeastern University), Amir Taherin (Northeastern University), Bin Ren (William & Mary), Yanzhi Wang (Northeastern University), Jennifer Dy (Northeastern University), David Kaeli (Northeastern University)
Lunch
WfPerf: Characterizing and Benchmarking Smartly for Distributed and In Situ HPC/AI Workflows
Hao Qi (University of Florida), Xiaoyi Lu (University of Florida)
NIXT: A NCCL Inspector Exporter Tool for Observability of Collective Communication in Large Model Training
Ziyang Jia (University of California, Riverside), Sirshak Das (Nvidia), Jason Sewall (Nvidia), Laxmi Bhuyan (University of California, Riverside), Pasha Shamis (Nvidia), Daniel Wong (University of California, Riverside)
ChainForge: Characterizing Embedding as the Bottleneck in Quantum Annealer Workloads
Kanishka Jayathilake (University of Colorado Boulder), Cordelia Brumley (University of Colorado Boulder), Tanner Smith (University of Colorado Boulder), Ramin Ayanzadeh (University of Colorado Boulder)
Architecting Instruction Execution for Scalable Fault-Tolerant Quantum Computing
Dmitry Filippov (University of Cambridge), Justin Hogaboam (Microsoft), Mathias Soeken (Microsoft), Raghu Gatta (Microsoft), Prakash Murali (University of Cambridge), Matthias Troyer (Microsoft)
Diphda: Demystifying Processing-in-HBM for SpGEMM across Diverse Data Mapping
Helya Hosseini (University of Maryland, College Park), Sanjali Yadav (University of Maryland), Christina Giannoula (Max Planck Institute for Software Systems (MPI-SWS)), Bahar Asgari (University of Maryland, College Park)
Characterizing Energy Efficiency Trade-offs in the Linux Storage Stack for Flash-based NVMe SSDs
Joseph Subash Kanichai (Vrije Universiteit Amsterdam), Krijn Doekemeijer (Vrije Universiteit Amsterdam), Steven van der Vlugt (ASTRON), Dante Niewenhuis (Vrije Universiteit Amsterdam), Tiziano de Matteis (Vrije Universiteit Amsterdam), Animesh Trivedi (IBM Research Europe, Zurich)
Characterizing Performance Bottleneck of Distributed Shared Memory in Modern GPUs
Jong Hyun Jeong (Korea University), Myung Kuk Yoon (Ewha Womans University), Yunho Oh (Korea University), Hyeran Jeon (University of California Merced), Gunjae Koo (Korea University)
MultiFACE: Scaling Out Analytical Soft-Error Vulnerability Evaluation for Multicore Architectures
George-Marios Fragkoulis (University of Athens), Dimitris Gizopoulos (University of Athens)
Break
PageSelect: A System-Level Approach for Selective Use of Page Coloring and Huge Pages
Theodoros Moraitis (University of Athens), Konstantinos Kordolaimis (University of Athens), Dimitris Gizopoulos (University of Athens), Vasileios Karakostas (University of Athens)
RapidQ: Queueing-based Performance Modeling Framework for Rapid Simulation and Automated Tuning of Input-Dependent Streaming FPGA Pipelines
Shashank Obla (Carnegie Mellon University), Bin Li (Intel), James C. Hoe (Carnegie Mellon University)
PANEM: A Heuristic Latency Model
Lucas Crowthers (Ampere Computing), Mahesh Madhav (Ampere Computing)
CXL-ClusterSim: Modeling CXL-based Disaggregated Memory Cluster for Pooling and Sharing using gem5 and SST
Kaustav Goswami (University of California, Davis), Maryam Babaie (University of California, Davis), Hoa Nguyen (University of California, Davis), Venkatesh Akella (University of California, Davis), Jason Lowe-Power (University of California, Davis)