PACT 2026   October 19–22, 2026

Sponsors

Platinum

Silver

Bronze

Program at a Glance

Monday, October 19, 2026

Time Program Discovery Room Classroom A Classroom B Classroom C
8:00–9:00 AM Registration Registration / Check-in
9:00–10:00 AM Workshops & Tutorials — Morning I Workshop: 1st National Science Data Fabric Summit.
Organizers: Michela Taufer, University of Tennessee, Knoxville; Valerio Pascucci, University of Utah.
Workshop: 1st Workshop on ML for Assisting Code Quality (MLAC).
Workshop Chair: Jay Lofstead.
Tutorial: CEDR: A Holistic Software and Hardware Design Environment for Hardware Agnostic Application Development and Deployment on FPGA-Integrated Heterogeneous Systems.
Presenters: Serhan Gener, Umut Suluhan, Ali Akoglu.
Tutorial: SODA Synthesizer: Accelerating Artificial Intelligence Applications with an End-to-End Silicon Compiler.
Presenters: Nicolas Bohm Agostini, Vito Giovanni Castellana, Fabrizio Ferrandi, Serena Curzel, Ankur Limaye, Antonino Tumeo.
10:00–10:30 AM Coffee Break Coffee Break
10:30 AM–12:00 PM Workshops & Tutorials — Morning II Workshop: 1st National Science Data Fabric Summit.
Organizers: Michela Taufer, University of Tennessee, Knoxville; Valerio Pascucci, University of Utah.
Workshop: 1st Workshop on ML for Assisting Code Quality (MLAC).
Workshop Chair: Jay Lofstead.
Tutorial: CEDR: A Holistic Software and Hardware Design Environment for Hardware Agnostic Application Development and Deployment on FPGA-Integrated Heterogeneous Systems.
Presenters: Serhan Gener, Umut Suluhan, Ali Akoglu.
Tutorial: SODA Synthesizer: Accelerating Artificial Intelligence Applications with an End-to-End Silicon Compiler.
Presenters: Nicolas Bohm Agostini, Vito Giovanni Castellana, Fabrizio Ferrandi, Serena Curzel, Ankur Limaye, Antonino Tumeo.
12:00–1:30 PM Lunch Lunch
1:30–3:00 PM Workshops & Tutorials — Afternoon I Workshop: 1st National Science Data Fabric Summit.
Organizers: Michela Taufer, University of Tennessee, Knoxville; Valerio Pascucci, University of Utah.
Workshop: 1st LACS (Learning-Augmented Compilers & Systems).
Organizers: Eun Jung (EJ) Park, Riyadh Baghdadi, Joseph Manzano, Keren Zhou.
Tutorial: Chameleon: An Open Platform for Computer Science Experimentation.
Presenters: Kate Keahey, Mark Powers.
Tutorial: Reproducible Benchmarking for High-Performance Computing Applications.
Presenters: Olga Pearce, Gregory Becker, Doug Jacobsen, Stephanie Brink
3:00–3:30 PM Coffee Break Coffee Break
3:30–5:00 PM Workshops & Tutorials — Afternoon II Workshop: 1st National Science Data Fabric Summit.
Organizers: Michela Taufer, University of Tennessee, Knoxville; Valerio Pascucci, University of Utah.
Workshop: 1st LACS (Learning-Augmented Compilers & Systems).
Organizers: Eun Jung (EJ) Park, Riyadh Baghdadi, Joseph Manzano, Keren Zhou.
Tutorial: Chameleon: An Open Platform for Computer Science Experimentation.
Presenters: Kate Keahey, Mark Powers.
Tutorial: Reproducible Benchmarking for High-Performance Computing Applications.
Presenters: Olga Pearce, Gregory Becker, Doug Jacobsen, Stephanie Brink
5:00–6:00 PM PACT 2026 Welcome Reception Supported by (NSDF)

Tuesday, October 20, 2026

Time Program Discovery Room
8:00–9:00 AM Registration Registration / Check-in
9:00–9:30 AM PACT 2026 Welcome & Opening Remarks Plenary — Discovery Room
9:30–10:30 AM Keynote — Andrew A. Chien Plenary — Discovery Room. Chair: Michela Taufer, University of Tennessee Knoxville
10:30–11:00 AM Coffee Break Coffee Break
11:00 AM–12:30 PM Best Paper Candidates — Plenary Session
Chair: Jaejin Lee, Seoul National University
11:00–11:30 AM #289: Hoppolyta: Polyhedral Kernel Generation Meets Hopper Architecture
Aravind Acharya; Somashekaracharya G Bhaskaracharya; Evghenii Gaburov; Bin Fan; Alexander Collins; Bastian Hagedorn; Vinod Grover
11:30 AM–12:00 PM #344: SALT: Symbolic Analysis of Loop Tiling
Yanghui Wu; Yifan Zhu; Yekai Pan; Chen Ding
12:00–12:30 PM #130: ESA: Improving GPU Utilization with Elastic Isolation for ML Inference Services
Taeklim Kim; Saurabh Agarwal; Rachata Ausavarungnirun; Jayneel Gandhi; Christopher J. Rossbach
12:30–2:00 PM Lunch Lunch
2:00–3:00 PM PACT 2026 Panel — Plenary Who Gets to Do Computing Research in 2036?
Panelists:
Eun Jung (EJ) Park — Qualcomm Innovation Center
Valerio Pascucci — University of Utah
Hariharan Devarajan — Lawrence Livermore National Laboratory
Tanu Malik — University of Missouri, Columbia
Lawrence Rauchwerger — UIUC
Moderator: Michela Taufer — University of Tennessee, Knoxville
3:00–3:30 PM Coffee Break Coffee Break
3:30–5:00 PM Poster Lightning Presentations 27 posters: 15 Research Posters + 12 ACM SRC Posters
Presentations up to 3 minutes each. Chair: Jay Lofstead, Sandia National Laboratories
5:00–6:00 PM PACT Poster & ACM SRC Reception Discovery Room

Wednesday, October 21, 2026

Time Program Discovery Room Classroom A Classroom C
8:00–9:00 AM Registration Registration / Check-in
9:00–10:00 AM Keynote — Josep Torrellas Plenary — Discovery Room. Chair: Jaejin Lee, Seoul National University
10:00–10:30 AM Coffee Break Coffee Break
10:30 AM–12:00 PM Technical Sessions 1–3 — Parallel Session 1 — Compiler Techniques for Dataflow and Accelerator Systems
Chair: To be announced
Session 2 — Processing-in-Memory for Large-Scale AI
Chair: To be announced
Session 3 — Compilation and Execution for Heterogeneous Architectures
Chair: To be announced
Talk 1 — #154: Stream Swizzling: Extending HLS for Efficient Non-Affine Stream Permutations
Chengyue Wang; Jason Cong; Jonathan Xue; Shinju Ju; Yingquan Wu
Talk 1 — #5: PIM-SHERPA: Software Method for On-device LLM Inference by Resolving PIM Memory Attribute and Layout Inconsistencies
Byeongho Kim; Hyeonsu Kim; Jaehoon Yu; Kyomin Sohn; Sanghoon Cha; Seungwon Lee; Seungwoo Seo; Sukhan Lee; Sunjung Lee; Yongjun Park; Yuhwan Ro
Talk 1 — #56: AIEHalide: Compiling Halide to Spatial NPU Dataflow with Constrained Autoscheduling
Abnikant Singh; Suresh Purini
Talk 2 — #188: Splyce: SIMD Vectorization of Sparse Coiteration
Kabilan Mahathevan; Kirshanthan Sundararajah; Poorna Gunathilaka
Talk 2 — #87: PDiMC: Achieving High-Throughput LLM Inference and Resolving Memory Concurrency via an Efficient PIM Subsystem
Byeongho Kim; Hweesoo Kim; Jaewan Choi; Kyomin Sohn; Sukhan Lee; Wontak Han; Yoonah Paik
Talk 2 — #127: FIFO Initialization: Efficient Support for Serial Loops on Spatial Elastic CGRAs
Eric Xu; Tarek Abdelrahman
Talk 3 — #240: ST-Flow: A Hardware Compiler for Automating Spatial-Temporal Dataflow Acceleration
Jason Cong; Stéphane Pouget; Suhail Basalama
Talk 3 — #109: A Hybrid Processing-in-Memory Architecture for Long Sequence LLM Inference with KV Cache Filtering
Jaehyuk Huh; Juhyun Lee; Sanghyeon Lee; Soojin Hwang
Talk 3 — #265: KERYX: A CUDA/HIP Framework for Adaptive Runtime Compilation in Heterogeneous Systems
Marc Gonzalez Tallada
Talk 4 — #409: Practical Correctness and Equivalence Checking for MLIR
Emily Tucker; Erika Hunhoff; Erwei Wang; Louis-Noel Pouchet; Stephen Neuendorffer
Talk 4 — #377: ASTRA-MoE: GPU-Augmenting In-Storage Acceleration for Long-Context Mixture-of-Expert Inference
Hyeonggyu Jeong; Inyoung Song; Jinwoo Jeong; Jungwook Choi; Kyungmo Koo; Yongho Song
Talk 4 — #364: Partial Instruction Execution on Long SIMD Architectures
Adrià Armejach; Francesc Martinez; Marc Casas
12:00–1:30 PM Lunch Lunch
1:30–3:00 PM Technical Sessions 4–6 — Parallel Session 4 — Memory Hierarchies and Processing-in-Memory
Chair: To be announced
Session 5 — Memory, Communication, and System-Level Data Movement
Chair: To be announced
Session 6 — Accelerating AI Across GPUs and Mobile Systems
Chair: To be announced
Talk 1 — #34: DREAM: In-DRAM Bit-Serial PIM with Data Reuse and Efficient Mapping
Aman Arora; Jeeho Ryoo; Jiajun Hu; Lizy K. John; Siddhartha Raman Sundara Raman; Siyuan Ma
Talk 1 — #318: Hermes: Accelerating Page Migration and HPC Data Transfers with NoC-attached Engines
Adrià Armejach; Francesco Sgherzi; Ivan Vega; Jordi Fornt; Juan Miguel de Haro Ruiz; Marco Siracusa; Miquel Moreto; Pouya Esmaili Dokht
Talk 1 — #195: Efficient Scheduling Algorithm for Large-scale Models on Heterogeneous Mobile Systems
Jinyoung Kim; Minseong Kim; Yongjun Park; Yongjun Yongjun
Talk 2 — #73: SPARQ: Skew-Aware PIM Accelerator for Relational Join and Select Queries
Sabiha Tajdari; Anastasia Ailamaki; Sandhya Dwarkadas
Talk 2 — #91: Cordelia: A Huffmanized Merkle Tree for Secure Memory
Galy Sela; Iris Bahar; Maurice Herlihy; Samuel Thomas; Tali Moreshet
Talk 2 — #253: PALRAC: Parallel Linear Recurrence Accelerator for Tree-based Speculative Decoding
Hyuk-Jae Lee; Sangheon Lee; Xuan Truong Nguyen
Talk 3 — #489: Focus on What Matters: DRAM-CXL Hybrid Memory Management with PRISM
Daniel Mosse; Fatemeh Golshan
Talk 3 — #143: A Coordinated Approach to Transactional Data Structures
Ahmed Hassan; Michael Spear; Yaodong Sheng
Talk 3 — #26: Split-Posit Systolic Array: A Resource-Efficient Hardware Accelerator for High-Performance AI Workloads
Arun M; Madhav Rao; Sneha Dandekar; Vaishnavi Sharma
Talk 4 — #548: ElaCache: Fine-Grain Dynamic Partitioning of LLCs and Coherence Directories in Multiprocessors
Adam Morrison; Dingyuan Cao; Josep Torrellas; Neil Zhao
Talk 4 — #21: Proba: A High-Performance, Low-Traffic Probabilistic Spatial Memory Streaming Prefetcher
Jacky Wong; Sam Ainsworth; Yinting Huang
Talk 4 — #107: HiNa-MoE: High-Performance, Non-Intrusive MoE Inference on CPUs with Matrix Engines
Dezun Dong; Enda Yu; Jianbin Fang; Junwen Zhang; Weiling Yang; Zhe Bai
3:00–3:30 PM Coffee Break Coffee Break
3:30–4:15 PM Technical Sessions 7–8 — Parallel Session 7 — Compilation for Specialized Computing
Chair: To be announced
Session 8 — Efficient ML Serving and Heterogeneous Scheduling
Chair: To be announced
ACM SRC Poster Finalists
Chair: Jay Lofstead, Sandia National Laboratories

Presenters will be announced on Tuesday, October 20, following the PACT Poster Session and Poster Reception.

Finalists for the ACM Student Research Competition (SRC) will be selected based on the quality of their poster presentations and discussions during the Poster Session. Selected finalists will be invited to present their work in this session.
Talk 1 — #277: Reducing Address Arithmetic Overheads: New Compiler Techniques for Programmable Dataflow-based AI Accelerators
Alberto Mannari; Alex Gatea; Bardia Mahjour; Chris Bowler; Masoud Ataei Jaliseh; Nicole Khoun; Prasanth Chatarasi; Shubham Jain; Swagath Venkataramani; Viji Srinivasan; Wei Wang
Talk 1 — #519: HeteroSched: Co-Optimizing Scheduling and Parallelization for Deep Learning Workloads for Heterogeneous GPU Clusters
Amirali Mirian; Bahram Afsharmanesh; Gagan Agrawal; Md Musfiqur Rahman Sanim
Talk 2 — #58: Q-TranSim: Batch Quantum Circuit Simulation using Tensor Transpilation
Hengrui Chen; Shui Jiang; Tsung-Wei Huang; Tsung-Yi Ho
Talk 2 — #633: EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models
Gagan Agrawal; Hangyu Zheng; Kunxiong Zhu; Miao Yin; Minghai Qin; Wei Niu; Zhihao Shu
4:15–5:00 PM Industry Keynote — Samantika Sury, HPE Plenary — Discovery Room. Chair: Antonino Tumeo, PNNL
6:30–10:00 PM PACT 2026 Banquet — Skydeck Chicago
PACT Awards presented during the banquet

Thursday, October 22, 2026

Time Program Discovery Room Classroom A Classroom C
8:00–9:00 AM Registration Registration / Check-in
9:00–10:00 AM Keynote — Mary Hall Plenary — Discovery Room. Chair: Tanu Malik, University of Missouri, Columbia
10:00–10:30 AM Coffee Break Coffee Break
10:30 AM–12:00 PM Technical Sessions 9–11 — Parallel Session 9 — GPU Execution and AI Workload Optimization
Chair: To be announced
Session 10 — GPU Memory Systems and Data-Intensive Acceleration
Chair: To be announced
Session 11 — Performance Optimization, Scheduling, and Communication
Chair: To be announced
Talk 1 — #141: Improving Data Reuse across Blocks for Efficient Block-sparse Transformers on GPUs
Lihan Hu; Peng Jiang; Sun Xian-He; Xian-He Sun; Xiaoyang Lu
Talk 1 — #39: SLGS: A Structure-Aware Scanline Renderer for Efficient 3D Gaussian Splatting
Dongho Ha; Hyunwuk Lee; Mingu Jung; Seunghyun Lee; Sungbin Kim; Sungwoo Kim; Won Woo Ro; Yingyan (Celine) Lin
Talk 1 — #97: Versatile Power Benchmark Generation for Modeling Emerging GPU Workloads and their Power Events
Allison Seigler; Lizy K. John; Zhixing Jiang
Talk 2 — #75: Fast Cross-Operator Optimization of Attention Dataflow on Spatial Accelerators
Bo Yuan; Hailiang Hu; Haodong Chang; Jiang Hu; Rongjian Liang; Yu Gong; Zhenrui Wang; Zhexiang Tang
Talk 2 — #191: GZswap: Hardware-Managed Compressed Swap for GPU Memory Oversubscription
Boyeol Choi; Jungrae Kim; Sanghyun Hong; Sanghyun Park; Seokin Hong
Talk 2 — #475: HSF: A Hierarchical Scheduling Framework for Application-Tailored Scheduling on Tasking Runtimes
Antoni Navarro Muñoz; David Álvarez; Vicenç Beltran; Vincent A. Arcila Larrea
Talk 3 — #327: KerColle: Unlocking Fine-Grained GPU Concurrency in Vision-Language-Action Models
Anna Li; Christina Giannoula; Nandita Vijaykumar
Talk 3 — #13: V-PWC: Accelerating Page Table Walk for Multi-Chip-Module GPUs
Mingxin Tang; Qiong Li; Sen Yang; Wei Chen; Xia Zhao; Xinjin Gao
Talk 3 — #146: Scalable Asynchronous Aggregation for Many-to-Many Transfers
Naveen Namashivayam Ravichandrasekaran; Pen-Chung Yew
Talk 4 — #405: Mask-Aware Execution for Efficient JEPA Training
Amirali Mirian; Bahram Afsharmanesh; Gagan Agrawal; Md Musfiqur Rahman Sanim; Wei Niu; Zhihao Shu
Talk 4 — #238: CAPO: Enhancing CXL-based Memory Expanders with Adaptive Prefetching
Anwen Huang; Chenglong Li; Jingyan Song; Na Nie; Qiong Li; Yongji Liu; Zhijie Liu
Talk 4 — #190: VaBO: Delivering Reliable Autotuning Performance in Parallel Applications under Unreliable Conditions
Hengrui Luo; Saeyeon Kim; Tirthak Patel; Vedica Rao; Younghyun Cho
12:00–12:10 PM Transition to Discovery Room — Please move to Discovery Room for the conference closing
12:10–12:20 PM PACT 2026 Closing Remarks Discovery Room

Program notes: All times are Central Time (CT). Session chairs and the poster-reception location will be updated when confirmed. Schedule subject to change.


Keynotes

Andrew A Chien Tuesday, October 20, 2026

UpDown: A Supercomputer co-designed for Graph Computing and Data Transformation

Andrew A Chien

Univ of Chicago and Argonne National Lab
Chicago UpDown Computing, Inc.

Supercomputers and AI compute are ill-suited for sparse and data-intensive computations because they are optimized for maximum dense matrix "FLOPS". We have designed the UpDown System – optimized for graph computing, streaming data ingestion and transformation as well as high-level programming. The result outperforms conventional CPU/GPU-based systems by 10-100x on an ISO-power basis.

UpDown's radical micro-architecture unleashes fine-grained parallelism: 1-cycle thread creation and management, 1-cycle messages. This enables efficient computation on 10-instruction thread invocations. Further, software-controlled split-transaction DRAM access unlocks the power of HBM's massive memory bandwidth. For irregular applications, UpDown datapath efficiency is 10x greater. UpDown performance on skewed-graph computations exceeds multicore CPU's (>100x) and GPU's (20-60x) in single-node configuration. Updown performance scales to 1,000 and 10,000-fold speedup on BFS, Pagerank, Triangle Count, K-truss and more.

UpDown data ingestion exceeds 5 billion/records/s/node (1000x CPU-based databases), reaching 100 trillion records/s. It enables a new class of streaming analytics and complex workflows. UpDown enables high level programming with a global address space, and a flexible map-reduce framework (KVMSR) coupled with an event-driven language (UDWeave). This enables easy vertex, edge-centric programming, and fits well for other data-parallel models such as relational/graphDB, sparse matrixes, and more.

UpDown was created under funding from IARPA's AGILE program, and is being commercialized by Chicago UpDown Computing, Inc. (www.chupdown.com).

Bio

Andrew A Chien is the William Eckhardt Distinguished Service Professor of Computer Science at the University of Chicago and Senior Scientist at Argonne National Laboratories. Chien led the IARPA funded "UpDown System Project", designing breakthrough scalable graph analytics systems and is now Founder and President of Chicago UpDown Computing, Inc. (www.chupdown.com). He has led the Zero-carbon Cloud project since 2015, and is known for his research on datacenters, renewable energy and sustainability, cloud resource management and software, and large-scale system architecture. Chien has received numerous recognitions for research. Dr. Chien currently serves on the NSF CISE Advisory Committee and DARPA ISAT. He is a Fellow of the ACM, IEEE, and AAAS. He served as EiC of Communications of the ACM, 2017-2022, and Vice President of Research at Intel Corporation from 2005-2010. He served as SAIC Chair Professor of University of California, San Diego (1998-2005) and as faculty at the University of Illinois (1990-98). He received BS, MS, and PhD degrees from the Massachusetts Institute of Technology.

Tuesday, October 20, 2026

PACT 2026 Panel

Who Gets to Do Computing Research in 2036?

Panelists: Eun Jung (EJ) Park — Qualcomm Innovation Center
Valerio Pascucci — University of Utah
Hariharan Devarajan — Lawrence Livermore National Laboratory
Tanu Malik — University of Missouri, Columbia
Lawrence Rauchwerger — UIUC
Moderator: Michela Taufer — University of Tennessee, Knoxville

Computing research is becoming increasingly expensive, complex, and concentrated. Emerging areas such as artificial intelligence and quantum computing often require specialized infrastructure, large datasets, substantial funding, technical staff, and extensive institutional capacity. Ambitious national initiatives such as the Genesis Mission further demonstrate the growing importance of coordinated research across academia, industry, national laboratories, and government. Yet the resources needed to participate in such efforts remain unevenly distributed.

Looking toward 2036, where will computing research take place, and who will be able to participate? Will the most consequential research become concentrated within a small number of well-resourced universities, companies, and national laboratories? What roles will smaller academic institutions, emerging companies, and individual researchers play?

This panel will bring together perspectives from academia, industry, and national laboratories to examine what counts as computing research, whether funding and infrastructure have become proxies for research excellence, and how institutional resources shape who can contribute. Panelists will discuss how cross-sector partnerships and national initiatives such as Genesis can broaden participation while preserving pathways for small teams, foundational and exploratory work, undergraduate-driven research, and institutions serving diverse students and regions.

Ultimately, the panel asks: How can we ensure that computing research in 2036 is shaped by the breadth and quality of its ideas—not only by where researchers work or the resources available to them?

Josep Torrellas Wednesday, October 21, 2026

Toward Accelerator-Centric Computing

Josep Torrellas

Thomas M. Siebel Chair in Computer Science
Director, SRC JUMP 2.0 ACE Center for Evolvable Computing
University of Illinois, Urbana-Champaign
iacoma.cs.uiuc.edu/josep/torrellas.html

Given current energy-consumption trends, there is ample consensus that we will have to move much of the computation to hardware accelerators. This is because accelerators are the most energy-efficient platforms. However, from the evidence of past efforts in this direction, architecting an accelerator-centric computing environment looks very challenging. It is unclear what architectural designs and software advances will really enable this new paradigm. In this talk, I will outline our vision of the hardware and software needed for a successful accelerator-centric computing environment, and some of the efforts that we are doing in this direction.

Bio

Josep Torrellas is the Thomas M. Siebel Chair in Computer Science at the University of Illinois, Urbana-Champaign (UIUC). He is the Director of the ACE Center for Evolvable Computing (an SRC/DARPA JUMP 2.0 Center), past Co-Leader of an Intel Strategic Research Alliance (ISRA) on Computer Security, and past Director of the Illinois-Intel Parallelism Center (I2PC). His research interests are multiprocessor computer architectures and parallel computing. Some of his contributions include thread-level speculation (TLS) architectures, the Bulk Multiprocessor concept, deterministic record and replay mechanisms, process variation mitigation techniques, and hardware defenses against speculative execution attacks. In addition, he has contributed to several experimental multiprocessor designs such as IBM's PERCS Multiprocessor, Intel's Runnemede Extreme-Scale Multiprocessor, Illinois Cedar, and Stanford DASH.

Torrellas has received the IEEE Computer Society (CS) Harry H. Goode Memorial Award, the UIUC Daniel C. Drucker Eminent Faculty Award, the UIUC Campus Award for Excellence in Graduate Student Mentoring, the IEEE CS Edward J. McCluskey Technical Achievement Award, and was a Willett Faculty Scholar at UIUC. He is an IEEE CS Golden Core Member, and a Fellow of IEEE, ACM, and AAAS. He was the Chair of the IEEE Technical Committee on Computer Architecture (TCCA). He has served in the Board of Directors of the Computing Research Association (CRA) and has been a Council Member of CRA's Computing Community Consortium (CCC). He was a member of the U.S. National Academies Board on Army Research and Development. He serves in the International Roadmap for Devices and Systems (IRDS). Torrellas has graduated 53 PhDs. He received a PhD from Stanford University.

Samantika Sury Wednesday, October 21, 2026

Tightly Coupled Customizability: Enabling the Next Generation of AI-HPC Systems

Samantika Sury

Fellow and Chief Hardware Architect
HPE - HPC and AI Infrastructure Solutions

The end of Moore's Law scaling and the rapid rise of AI are driving a fundamental shift in computer architecture, accelerating the adoption of purpose-built technologies across compute, memory, networking, and storage. At the same time, scientific computing is moving beyond isolated applications toward tightly integrated workflows that combine simulation, data analytics, learning, inference, and increasingly agentic forms of execution. Together, these changes are reshaping how systems are designed and where performance bottlenecks emerge.

Next-generation AI-HPC systems will depend on tightly coupled and customizable architectures that bring specialized resources together around the needs of complete workflows rather than individual applications. This keynote explores architectural directions including workflow-centric optimization, flexible scale-up and scale-out fabrics, and macroheterogeneity. It will examine the trends driving these changes, the challenges they introduce, and the opportunities they create as the community moves toward more integrated, adaptable, and workload-aware AI-HPC systems.

Bio

Samantika Sury serves as an HPE Fellow, Vice President, and Chief Hardware Architect for HPC and AI Infrastructure Solutions. She leads the Future Technologies team, which focuses on advancing hardware and software system innovations. Samantika has previously held prominent roles at Samsung, where she served as Vice President and Chief Hardware Architect for HPC. She has also worked at Intel® as a Senior Principal Engineer, driving silicon and system architecture innovations into marketable products, served as the Chief Architect of Intel's HPC-Custom Silicon Program and was the Principal Investigator and Lead Architect for the DOE PathForward Program. Samantika holds 31 U.S. and international patents, has published more than 20 peer-reviewed papers, and has delivered numerous invited talks. She was recognized in HPCWire People to Watch 2026. Samantika earned her Ph.D. in Computer Science from the Georgia Institute of Technology.

Mary Hall Thursday, October 22, 2026

Tiles, Bricks, and Layouts: How Aggregate Data Abstractions Aid in Optimizing Data Movement

Mary Hall

Professor, Kahlert School of Computing, University of Utah

Data movement is the dominant execution and energy cost across the application workloads in data centers and supercomputers. Programming at the tile level has become a popular strategy for optimizing data movement for both deep learning and general structured grids, using Triton, cuTile, bricks, and fine-grained data blocks. Expressing hierarchical data and thread layouts, mostly designed with matrix processors in mind, facilitates automatic code generation that further raises the level of abstraction in such code. In this talk, we will describe prior work on BrickLib supporting fine-grained data blocks and active research on LEGO for hierarchical data and thread layout. We will connect these concepts with emerging hardware features and future demands on programming systems to reduce data movement.

Bio

Mary Hall is a Professor and former Director of the Kahlert School of Computing at University of Utah. Her research focuses on high-performance computing, compiler optimizations and code generation for novel and emerging hardware, and performance tuning. She has served on the Board of Directors of the Computing Research Association since 2015, and she is currently its Vice Chair. She is an ACM Distinguished Scientist and an IEEE Fellow.

Papers

Title Authors
PIM-SHERPA: Software Method for On-device LLM Inference by Resolving PIM Memory Attribute and Layout Inconsistencies Sunjung Lee (Samsung Advanced Institute of Technology); Sanghoon Cha (Samsung Advanced Institute of Technology); Hyeonsu Kim (Samsung Advanced Institute of Technology); Seungwoo Seo (Samsung Advanced Institute of Technology); Yuhwan Ro (Samsung Advanced Institute of Technology); Sukhan Lee (Samsung Electronics); Byeongho Kim (Samsung Electronics); Yongjun Park (Yonsei University); Kyomin Sohn (Samsung Electronics); Seungwon Lee (Samsung Advanced Institute of Technology); Jaehoon Yu (Samsung Advanced Institute of Technology)
V-PWC: Accelerating Page Table Walk for Multi-Chip-Module GPUs Sen Yang (College of Computer Science and Technology, National University of Defense Technology); Wei Chen (College of Computer Science and Technology, National University of Defense Technology); Xinjin Gao (College of Computer Science and Technology, National University of Defense Technology); Mingxin Tang (College of Computer Science and Technology, National University of Defense Technology); Qiong Li (Defense Innovation Institute); Xia Zhao (Defense Innovation Institute)
Proba: A High-Performance, Low-Traffic Probabilistic Spatial Memory Streaming Prefetcher Yinting Huang (Huawei Newton Research Centre); Jacky Wong (Imperial College London); Sam Ainsworth (University of Edinburgh)
Split-Posit Systolic Array: A Resource-Efficient Hardware Accelerator for High-Performance AI Workloads Sneha Dandekar (IIIT Bangalore); Arun M (IIIT Bangalore); Vaishnavi Sharma (IIIT Bangalore); Madhav Rao (IIIT Bangalore)
DREAM: In-DRAM Bit-Serial PIM with Data Reuse and Efficient Mapping Siyuan Ma (University of Texas at Austin); Jiajun HU (Arizona State University); Jeeho Ryoo (Fairleigh Dickinson University); Siddhartha Raman Sundara Raman (The University of Texas at Austin); Aman Arora (Arizona State University); Lizy K. John (University of Texas at Austin)
SLGS: A Structure-Aware Scanline Renderer for Efficient 3D Gaussian Splatting seunghyun lee (Yonsei University); Mingu Jung (Yonsei University); Sungwoo Kim (Yonsei University); Sungbin Kim (Yonsei University); Dongho Ha (Meta Platforms); Hyunwuk Lee (Unaffiliated); Yingyan (Celine) Lin (Georgia Institute of Technology); Won Woo Ro (Yonsei University)
AIEHalide: Compiling Halide to Spatial NPU Dataflow with Constrained Autoscheduling Abnikant singh (IIIT Hyderabad); Abnikant Singh (AMD, Inc); Suresh Purini (IIIT Hyderabad)
Q-TranSim: Batch Quantum Circuit Simulation using Tensor Transpilation Shui Jiang (The Chinese University of Hong Kong); Hengrui Chen (Zhejiang University); Tsung-Yi Ho (The Chinese University of Hong Kong); Tsung-Wei huang (UW Madison)
SPARQ: Skew-Aware PIM Accelerator for Relational Join and Select Queries Sabiha Tajdari (University of Virginia); Anastasia Ailamaki (École Polytechnique Fédérale de Lausanne); Sandhya Dwarkadas (University of Virginia)
Fast Cross-Operator Optimization of Attention Dataflow on Spatial Accelerators Haodong Chang (Texas A&M University); Hailiang Hu (AMD Inc); Zhenrui Wang (Texas A&M University); Yu Gong (Amazon Web Services); Rongjian Liang (Nvidia); Zhexiang Tang (Rutgers University); Bo Yuan (Rutgers University); Jiang Hu (Texas A&M University)
PDiMC: Achieving High-Throughput LLM Inference and Resolving Memory Concurrency via an Efficient PIM Subsystem Byeongho Kim (Samsung Electronics); Kyomin Sohn (Samsung Electronics); Sukhan Lee (Samsung Electronics); Hweesoo Kim (Samsung Electronics); Wontak Han (Samsung Electronics); Yoonah Paik (Samsung Electronics); Jaewan Choi (Samsung Electronics)
Cordelia: A Huffmanized Merkle Tree for Secure Memory Samuel Thomas (Pomona College); Galy Sela (EPFL); Tali Moreshet (Boston University); Maurice Herlihy (Brown University); Iris Bahar (Colorado School of Mines)
Versatile Power Benchmark Generation for Modeling Emerging GPU Workloads and their Power Events Allison Seigler (University of Texas at Austin); Zhixing Jiang (University of Texas at Austin); Lizy K. John (University of Texas at Austin)
HiNa-MoE: High-Performance, Non-Intrusive MoE Inference on CPUs with Matrix Engines Weiling Yang (National University of Defense Technology); Junwen Zhang (National University of Defense Technology); Dezun Dong (National University of Defense Technology); Jianbin Fang (National University of Defense Technology); Enda Yu (National University of Defense Technology); Zhe Bai (National University of Defense Technology); Xiaopeng Deng (National University of Defense Technology)
A Hybrid Processing-in-Memory Architecture for Long Sequence LLM Inference with KV Cache Filtering Soojin Hwang (ETRI); Sanghyeon Lee (KAIST); Juhyun Lee (KAIST); Jaehyuk Huh (KAIST)
FIFO Initialization: Efficient Support for Serial Loops on Spatial Elastic CGRAs Eric Xu (University of Toronto); Tarek S. Abdelrahman (University of Toronto)
ESA: Improving GPU Utilization with Elastic Isolation for ML Inference Services Taeklim Kim (The University of Texas at Austin); Saurabh Agarwal (The University of Texas at Austin); Rachata Ausavarungnirun (MangoBoost Inc.); Jayneel Gandhi (Meta); Christopher J. Rossbach (UT Austin and Microsoft)
Improving Data Reuse across Blocks for Efficient Block-sparse Transformers on GPUs Lihan Hu (The University of Iowa); Xiaoyang Lu (Illinois Institute of Technology); Xian-He Sun (Illinois Institute of Technology); Peng Jiang (The University of Iowa)
A Coordinated Approach to Transactional Data Structures Yaodong Sheng (Lehigh University); Leoul Demissie (Lehigh University); Ahmed Hassan (Lehigh University); Michael Spear (Lehigh University)
Scalable Asynchronous Aggregation for Many-to-Many Transfers Naveen Namashivayam Ravichandrasekaran (University of Minnesota); Nathan Wichmann (Hewlett Packard Enterprise); Pen-Chung Yew (University of Minnesota)
Stream Swizzling: Extending HLS for Efficient Non-Affine Stream Permutations Chengyue Wang (UCLA); JONATHAN XUE (UCLA); LANCE GIANG (UCLA); SHINJU JU (UCLA); Yingquan Wu (MBZU AI Lab); Jason Cong (UCLA)
Splyce: SIMD Vectorization of Sparse Coiteration Kabilan Mahathevan (Virginia Tech); Poorna Gunathilaka (Virginia Tech); Kirshanthan Sundararajah (Virginia Tech)
VaBO: Delivering Reliable Autotuning Performance in Parallel Applications under Unreliable Conditions Vedica Rao (Santa Clara University); Saeyeon Kim (Santa Clara University); Hengrui Luo (Rice University); Tirthak Patel (Rice University); Younghyun Cho (Santa Clara University)
GZswap: Hardware-Managed Compressed Swap for GPU Memory Oversubscription Boyeol Choi (Sungkyunkwan University); Sanghyun Park (FuriosaAI); Sanghyun Hong (Sungkyunkwan University); Seokin Hong (Sungkyunkwan University); Jungrae Kim (Sungkyunkwan University)
Efficient Scheduling Algorithm for Large-scale Models on Heterogeneous Mobile Systems Jinyoung Kim (Yonsei University); Yongjun Kim (Samsung Electronics); Minseong Kim (Samsung Electronics); Yongjun Park (Yonsei University)
CAPO: Enhancing CXL-based Memory Expanders with Adaptive Prefetching Jingyan Song (Academy of Military Sciences); Chenglong Li (Academy of Military Sciences); Anwen Huang (Academy of Military Sciences); Qiong Li (Academy of Military Sciences); Yongji Liu (Academy of Military Sciences); Zhijie Liu (Academy of Military Sciences); Na Nie (Academy of Military Sciences)
ST-Flow: A Hardware Compiler for Automating Spatial-Temporal Dataflow Acceleration Suhail Basalama (UCLA); Stéphane Pouget (University of California, Los Angeles); Jason Cong (UCLA)
PALRAC: Parallel Linear Recurrence Accelerator for Tree-based Speculative Decoding Sangheon Lee (Seoul National University); Hyuk-Jae Lee (Seoul National University); Xuan Truong Nguyen (Seoul National University)
KERYX: A CUDA/HIP Framework for Adaptive Runtime Compilation in Heterogeneous Systems MARC GONZALEZ TALLADA (Universitat Politecnica de Catalunya); PEDRO VALERO (ORNL); KEITA TERANISHI; JEFF VETTER (ORNL)
Reducing Address Arithmetic Overheads: New Compiler Techniques for Programmable Dataflow-based AI Accelerators Prasanth Chatarasi (IBM Research); Nicole Khoun (IBM); Wei Wang (IBM); Chris Bowler (IBM); Alex Gatea (IBM); Shubham Jain (IBM Research); Masoud Ataei Jaliseh (IBM); Alberto Mannari (IBM); Bardia Mahjour (IBM); Viji Srinivasan (IBM (Research)); Swagath Venkataramani (IBM Research)
Hoppolyta: Polyhedral Kernel Generation Meets Hopper Architecture Aravind Acharya (NVIDIA); Somashekaracharya G Bhaskaracharya (NVIDIA); Evghenii Gaburov (NVIDIA); Bin Fan (NVIDIA); Alexander Collins (NVIDIA); Bastian Hagedorn (NVIDIA); Vinod Grover (NVIDIA)
Hermes: Accelerating Page Migration and HPC Data Transfers with NoC-attached engines Francesco Sgherzi (SiPearl, Barcelona Supercomputing Center); Juan Miguel de Haro Ruiz (Barcelona Supercomputing Center); Jordi Fornt (Barcelona Supercomputing Center); Pouya Esmaili Dokht (Barcelona Supercomputing Center); Marco Siracusa (Barcelona Supercomputing Center); Ivan Fernandez (Barcelona Supercomputing Center); Adrià Armejach (UPC/BSC); Miquel Moreto (UPC/BSC)
KerColle: Unlocking Fine-Grained GPU Concurrency in Vision-Language-Action Models Anna Li (University of Toronto); Christina Giannoula (Max Planck Institute for Software Systems (MPI-SWS)); Nandita Vijaykumar (University of Toronto)
SALT: Symbolic Analysis of Loop Tiling Yanghui Wu (University of Rochester); Yifan Zhu (University of Rochester); Yekai Pan (University of Rochester); Chen Ding (University of Rochester)
Partial Instruction Execution on Long SIMD Architectures Francesc Martinez (Barcelona Supercomputing Center); Adrià Armejach (UPC/BSC); Marc Casas (Barcelona Supercomputing Center)
ASTRA-MoE: GPU-Augmenting In-Storage Acceleration for Long-Context Mixture-of-Expert Inference Hyeonggyu Jeong (Hanyang University); Inyoung Song (Hanyang University); Jinwoo Jeong (Hanyang University); Kyungmo Koo (Hanyang University); Byungmin Ahn (Samsung Electronics); Dong-Min Shin (Samsung Electronics); Yong Ho Song (Samsung Electronics); Jungwook Choi (Hanyang University)
Mask-Aware Execution for Efficient JEPA Training Md Musfiqur Rahman Sanim (University of Georgia); Zhihao Shu (University of Georgia); Bahram Afsharmanesh (University of Georgia); Amirali Mirian (University of Georgia); Wei Niu (University of Georgia); Gagan Agrawal (University of Georgia)
Practical Correctness and Equivalence Checking for MLIR Emily Tucker (Georgia Institute of Technology); Louis-Noel Pouchet (Colorado State University); Erika Hunhoff (AMD); Stephen Neuendorffer (AMD); Erwei Wang (AMD)
HSF: A Hierarchical Scheduling Framework for Application-Tailored Scheduling on Tasking Runtimes Vincent A. Arcila Larrea (Barcelona Supercomputing Center); David Álvarez Robert (Barcelona Supercomputing Center); Antoni Navarro Muñoz (Barcelona Supercomputing Center); Vicenç Beltran (Barcelona Supercomputing Center)
Focus on What Matters: DRAM-CXL Hybrid Memory Management with PRISM Fatemeh Golshan (University of Pittsburgh); Daniel Mosse (University of Pittsburgh)
HeteroSched: Co-Optimizing Scheduling and Parallelization for Deep Learning Workloads for Heterogeneous GPU Clusters Bahram Afsharmanesh (University of Georgia); Md Musfiqur Rahman Sanim (University of Georgia); Amirali Mirian (University of Georgia); Gagan Agrawal (University of Georgia)
ElaCache: Fine-Grain Dynamic Partitioning of LLCs and Coherence Directories in Multiprocessors Dingyuan Cao (University of Illinois at Urbana-Champaign); Neil Zhao (UT Austin & NVIDIA); Adam Morrison (Tel Aviv University); Josep Torrellas (Univ. of Illinois Urbana-Champaign)
EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models Kunxiong Zhu (University of Georgia); Zhihao Shu (University of Georgia); Hangyu Zheng (University of Georgia); Minghai Qin (Western Digital Research); Miao Yin (University of Texas at Arlington); Gagan Agrawal (University of Georgia); Wei Niu (University of Georgia)

Research Posters

Title Authors (Affiliations)
SAIL: SRAM-Accelerated LLM Inference System with Lookup-Table-based GEMV Jingyao Zhang (University of California, Riverside); Jaewoo Park (Ulsan National Institute of Science and Technology); Jongeun Lee (Ulsan National Institute of Science and Technology); Elaheh Sadredini (University of California, Riverside)
BenchCPU: Performance Distribution-Aware CPU Benchmarking over Open Configuration Spaces Chenxi Wang (Institute of Computing Technology, Chinese Academy of Sciences); Yuchen Su (State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences); Lei Wang (Institute of Computing Technology, Chinese Academy of Sciences); Guoxin Kang (University of Chinese Academy of Sciences); Wanling Gao (Institute of Computing Technology, Chinese Academy of Sciences); Fan Zhang (The International Open Benchmark Council); Jianfeng Zhan (Institute of Computing Technology, Chinese Academy of Sciences)
Dynamic Scheduling of LLM Inference in Asymmetric Memory Systems for Energy-Performance Balance Soojin Hwang (ETRI); Jungwoo Kim (Stanford University); Sanghyeon Lee (KAIST); Hongbeen Kim (KAIST); Jaehyuk Huh (KAIST)
ABS: Efficiently Serving Deep Learning Models with Offline CVMs Bojian Zheng (University of Toronto, CentML, Vector Institute); Wenyang Wang (University of Toronto, CentML, Vector Institute); Hui Shi (Tencent); Jie Fu (Tencent); Xiaofeng Yang (Tencent); Yangyu Tao (Tencent); Peng Chen (Tencent); Jie Jiang (Tencent)
FineCAT: Instruction-Count-Driven Fine-Grained LLC Management for Co-Location Scenarios Yanqi Kan (Institute of Computing Technology, Chinese Academy of Sciences); Lei Wang (Institute of Computing Technology, Chinese Academy of Sciences); Fanda Fan (University of Chinese Academy of Sciences); Yikang Yang (Institute of Computing Technology, Chinese Academy of Sciences); Wanling Gao (Institute of Computing Technology, Chinese Academy of Sciences); Chunjie Luo (Institute of Computing Technology, Chinese Academy of Sciences); Jianfeng Zhan (Institute of Computing Technology, Chinese Academy of Sciences)
HARMONY: A Framework for Cooperative Memory Scheduling in CXL Memory Systems Yongho Lee (Sungkyunkwan University); Junbum Park (Sungkyunkwan University); Sungbin Jang (Sungkyunkwan University); Osang Kwon (Samsung Electronics); Minkyu Choi (Samsung Electronics); Seokin Hong (Sungkyunkwan University)
HERO: Local HBM Enhancements to Support Remote Memory Optimizations Christin David Bose (Purdue University); Cesar Avalos (Purdue University); Yechen Liu (Purdue University); Timothy Rogers (Purdue University)
Automating Exploration and Code Generation for Pipelined Dataflows on AMD XDNA™ NPUs Joren Dumoulin (KU Leuven); Arne Symons (KU Leuven); Erika Hunhoff (AMD); Andre Roesti (AMD); Andra Bisca (AMD); Gagandeep Singh (AMD); Kristof Denolf (AMD); Marian Verhelst (KU Leuven)
TPE: AI-MicroBMT: A Unified Benchmark Tool for Deployment-Aware Performance Characterization of Diverse AI Accelerators Jonghyun Shin (Seoul National University); Dongmyong Shin (Seoul National University); Jeongnam Youn (Suwon Science College); Dae-Hwan Kim (Seoul National University); Soojung Ryu (Seoul National University); Xuan Truong Nguyen (Seoul National University); Hyuk-Jae Lee (Seoul National University)
Diffusion-Based Data Augmentation for Multi-Label Performance Modeling Mohammad Ali (Texas State University); Apan Qasem (Texas State University)
UniFlow: A Spatial Transformer Accelerator with Unified GEMM and Non-GEMM Dataflow Architecture Haocheng Xu (University of California, Irvine); Faraz Tahmasebi (University of California, Irvine); Rachid Karami (University of California, Irvine); Zhiheng Chen (University of California, Irvine); Hyoukjun Kwon (University of California, Irvine); Sitao Huang (University of California, Irvine)
NDS: Programmer-Free Offload of High-Performance Near Data Strands Shreyas Singh (University of Utah); Pratyush Nandi (University of Utah); Lin Jia (Intel); Shankar Balachandran (University of Utah); Rajeev Balasubramonian (University of Utah)
Cooperative Wavefronts: WFA and GWFA on a 4096-PE MIMD Many-Core Wenjie Geng (University of Michigan - Ann Arbor); Noah Kaplan (University of Michigan - Ann Arbor); Reetuparna Das (University of Michigan - Ann Arbor); Nathaniel Bleier (University of Michigan - Ann Arbor)
SCISSOR: Scalable I/O for Small Scattered Objects Runtime Kevin Assogba (Rochester Institute of Technology); Nigel Tan (Los Alamos National Laboratory); M. Mustafa Rafique (Rochester Institute of Technology); Michela Taufer (University of Tennessee Knoxville); Bogdan Nicolae (Argonne National Laboratory)
From Resampling to Retrieval: A Unified Budgeted View of Training with Large Candidate Sets Ziyang Jia (University of Missouri, Columbia); Tanu Malik (University of Missouri, Columbia)

ACM SRC

Title Authors (Affiliations)
RoadBlock: Rethinking GPU Tensor Core Microarchitecture for Emerging Microscaling Format Support Nikhil Rout (University of California, Los Angeles, United States)
Diagnosing GPU Performance Regressions Across Compiler Versions with Strata Befikir Bogale (University of Tennessee, United States)
Software-Controlled GB-Scale 3D-SRAM Residency for GPUs Eric Dubberstein (Carnegie Mellon University, United States)
Causal Observability for Microarchitectural Analysis Saber Ganjisaffar (University of California, Riverside, United States)
BIMBA: Best-Effort In-Network Merging with Bank-Side Accumulation for Transformer Accelerators Jungwoo Park (Seoul National University of Science and Technology, South Korea)
SpeedSparse: Exploiting Dense Matrix Multiplication for Practical Sparse Neural Networks Shreya Alladi (Computer Engineering Department, University of Murcia, Spain)
Model-Scale-Dependent Effects of Thread-Count Scaling on INT8 Dynamic Quantization Nahla Nabil Skaik (Arab Open University - Bahrain, Bahrain)
ACIO: Always-Complete Isolation of Outliers with Parallelism-Amortized Hardware for Low-Bit LLM Quantization Jihyeon Hwang (Seoul National University of Science and Technology, South Korea)
Joint Bit-Width and Fan-In Sensitivity Optimization for FHE-Aware Neural Network Acceleration Murat Toprak (Istanbul Technical University, Turkey)
Roofline-Decomposed Agents for Sample-Efficient On-Device LLM Execution Kaiyuan Zhang (Univeristy of Georgia, United States)
XBM: Hybrid Stack Composition for 3D Memory-on-GPU Inference Atharva Raut (Carnegie Mellon University, United States)