| Time | Program | Discovery Room | Classroom A | Classroom B | Classroom C |
|---|---|---|---|---|---|
| 8:00–9:00 AM | Registration | Registration / Check-in | |||
| 9:00–10:00 AM | Workshops & Tutorials — Morning I | Workshop: 1st National Science Data Fabric Summit. Organizers: Michela Taufer, University of Tennessee, Knoxville; Valerio Pascucci, University of Utah. |
Workshop: 1st Workshop on ML for Assisting Code Quality (MLAC). Workshop Chair: Jay Lofstead. |
Tutorial: CEDR: A Holistic Software and Hardware Design Environment for Hardware Agnostic Application Development and Deployment on FPGA-Integrated Heterogeneous Systems. Presenters: Serhan Gener, Umut Suluhan, Ali Akoglu. |
Tutorial: SODA Synthesizer: Accelerating Artificial Intelligence Applications with an End-to-End Silicon Compiler. Presenters: Nicolas Bohm Agostini, Vito Giovanni Castellana, Fabrizio Ferrandi, Serena Curzel, Ankur Limaye, Antonino Tumeo. |
| 10:00–10:30 AM | Coffee Break | Coffee Break | |||
| 10:30 AM–12:00 PM | Workshops & Tutorials — Morning II | Workshop: 1st National Science Data Fabric Summit. Organizers: Michela Taufer, University of Tennessee, Knoxville; Valerio Pascucci, University of Utah. |
Workshop: 1st Workshop on ML for Assisting Code Quality (MLAC). Workshop Chair: Jay Lofstead. |
Tutorial: CEDR: A Holistic Software and Hardware Design Environment for Hardware Agnostic Application Development and Deployment on FPGA-Integrated Heterogeneous Systems. Presenters: Serhan Gener, Umut Suluhan, Ali Akoglu. |
Tutorial: SODA Synthesizer: Accelerating Artificial Intelligence Applications with an End-to-End Silicon Compiler. Presenters: Nicolas Bohm Agostini, Vito Giovanni Castellana, Fabrizio Ferrandi, Serena Curzel, Ankur Limaye, Antonino Tumeo. |
| 12:00–1:30 PM | Lunch | Lunch | |||
| 1:30–3:00 PM | Workshops & Tutorials — Afternoon I | Workshop: 1st National Science Data Fabric Summit. Organizers: Michela Taufer, University of Tennessee, Knoxville; Valerio Pascucci, University of Utah. |
Workshop: 1st LACS (Learning-Augmented Compilers & Systems). Organizers: Eun Jung (EJ) Park, Riyadh Baghdadi, Joseph Manzano, Keren Zhou. |
Tutorial: Chameleon: An Open Platform for Computer Science Experimentation. Presenters: Kate Keahey, Mark Powers. |
Tutorial: Reproducible Benchmarking for High-Performance Computing Applications. Presenters: Olga Pearce, Gregory Becker, Doug Jacobsen, Stephanie Brink |
| 3:00–3:30 PM | Coffee Break | Coffee Break | |||
| 3:30–5:00 PM | Workshops & Tutorials — Afternoon II | Workshop: 1st National Science Data Fabric Summit. Organizers: Michela Taufer, University of Tennessee, Knoxville; Valerio Pascucci, University of Utah. |
Workshop: 1st LACS (Learning-Augmented Compilers & Systems). Organizers: Eun Jung (EJ) Park, Riyadh Baghdadi, Joseph Manzano, Keren Zhou. |
Tutorial: Chameleon: An Open Platform for Computer Science Experimentation. Presenters: Kate Keahey, Mark Powers. |
Tutorial: Reproducible Benchmarking for High-Performance Computing Applications. Presenters: Olga Pearce, Gregory Becker, Doug Jacobsen, Stephanie Brink |
| 5:00–6:00 PM | PACT 2026 Welcome Reception | Supported by (NSDF) | |||
| Time | Program | Discovery Room |
|---|---|---|
| 8:00–9:00 AM | Registration | Registration / Check-in |
| 9:00–9:30 AM | PACT 2026 Welcome & Opening Remarks | Plenary — Discovery Room |
| 9:30–10:30 AM | Keynote — Andrew A. Chien | Plenary — Discovery Room. Chair: Michela Taufer, University of Tennessee Knoxville |
| 10:30–11:00 AM | Coffee Break | Coffee Break |
| 11:00 AM–12:30 PM | Best Paper Candidates — Plenary Session Chair: Jaejin Lee, Seoul National University |
|
| 11:00–11:30 AM | #289: Hoppolyta: Polyhedral Kernel Generation Meets Hopper Architecture Aravind Acharya; Somashekaracharya G Bhaskaracharya; Evghenii Gaburov; Bin Fan; Alexander Collins; Bastian Hagedorn; Vinod Grover |
|
| 11:30 AM–12:00 PM | #344: SALT: Symbolic Analysis of Loop Tiling Yanghui Wu; Yifan Zhu; Yekai Pan; Chen Ding |
|
| 12:00–12:30 PM | #130: ESA: Improving GPU Utilization with Elastic Isolation for ML Inference Services Taeklim Kim; Saurabh Agarwal; Rachata Ausavarungnirun; Jayneel Gandhi; Christopher J. Rossbach |
|
| 12:30–2:00 PM | Lunch | Lunch |
| 2:00–3:00 PM | PACT 2026 Panel — Plenary | Who Gets to Do Computing Research in 2036? Panelists: Eun Jung (EJ) Park — Qualcomm Innovation Center Valerio Pascucci — University of Utah Hariharan Devarajan — Lawrence Livermore National Laboratory Tanu Malik — University of Missouri, Columbia Lawrence Rauchwerger — UIUC Moderator: Michela Taufer — University of Tennessee, Knoxville |
| 3:00–3:30 PM | Coffee Break | Coffee Break |
| 3:30–5:00 PM | Poster Lightning Presentations | 27 posters: 15 Research Posters + 12 ACM SRC Posters Presentations up to 3 minutes each. Chair: Jay Lofstead, Sandia National Laboratories |
| 5:00–6:00 PM | PACT Poster & ACM SRC Reception | Discovery Room |
| Time | Program | Discovery Room | Classroom A | Classroom C |
|---|---|---|---|---|
| 8:00–9:00 AM | Registration | Registration / Check-in | ||
| 9:00–10:00 AM | Keynote — Josep Torrellas | Plenary — Discovery Room. Chair: Jaejin Lee, Seoul National University | ||
| 10:00–10:30 AM | Coffee Break | Coffee Break | ||
| 10:30 AM–12:00 PM | Technical Sessions 1–3 — Parallel | Session 1 — Compiler Techniques for Dataflow and Accelerator Systems Chair: To be announced |
Session 2 — Processing-in-Memory for Large-Scale AI Chair: To be announced |
Session 3 — Compilation and Execution for Heterogeneous Architectures Chair: To be announced |
| Talk 1 — #154: Stream Swizzling: Extending HLS for Efficient Non-Affine Stream Permutations Chengyue Wang; Jason Cong; Jonathan Xue; Shinju Ju; Yingquan Wu |
Talk 1 — #5: PIM-SHERPA: Software Method for On-device LLM Inference by Resolving PIM Memory Attribute and Layout Inconsistencies Byeongho Kim; Hyeonsu Kim; Jaehoon Yu; Kyomin Sohn; Sanghoon Cha; Seungwon Lee; Seungwoo Seo; Sukhan Lee; Sunjung Lee; Yongjun Park; Yuhwan Ro |
Talk 1 — #56: AIEHalide: Compiling Halide to Spatial NPU Dataflow with Constrained Autoscheduling Abnikant Singh; Suresh Purini |
||
| Talk 2 — #188: Splyce: SIMD Vectorization of Sparse Coiteration Kabilan Mahathevan; Kirshanthan Sundararajah; Poorna Gunathilaka |
Talk 2 — #87: PDiMC: Achieving High-Throughput LLM Inference and Resolving Memory Concurrency via an Efficient PIM Subsystem Byeongho Kim; Hweesoo Kim; Jaewan Choi; Kyomin Sohn; Sukhan Lee; Wontak Han; Yoonah Paik |
Talk 2 — #127: FIFO Initialization: Efficient Support for Serial Loops on Spatial Elastic CGRAs Eric Xu; Tarek Abdelrahman |
||
| Talk 3 — #240: ST-Flow: A Hardware Compiler for Automating Spatial-Temporal Dataflow Acceleration Jason Cong; Stéphane Pouget; Suhail Basalama |
Talk 3 — #109: A Hybrid Processing-in-Memory Architecture for Long Sequence LLM Inference with KV Cache Filtering Jaehyuk Huh; Juhyun Lee; Sanghyeon Lee; Soojin Hwang |
Talk 3 — #265: KERYX: A CUDA/HIP Framework for Adaptive Runtime Compilation in Heterogeneous Systems Marc Gonzalez Tallada |
||
| Talk 4 — #409: Practical Correctness and Equivalence Checking for MLIR Emily Tucker; Erika Hunhoff; Erwei Wang; Louis-Noel Pouchet; Stephen Neuendorffer |
Talk 4 — #377: ASTRA-MoE: GPU-Augmenting In-Storage Acceleration for Long-Context Mixture-of-Expert Inference Hyeonggyu Jeong; Inyoung Song; Jinwoo Jeong; Jungwook Choi; Kyungmo Koo; Yongho Song |
Talk 4 — #364: Partial Instruction Execution on Long SIMD Architectures Adrià Armejach; Francesc Martinez; Marc Casas |
||
| 12:00–1:30 PM | Lunch | Lunch | ||
| 1:30–3:00 PM | Technical Sessions 4–6 — Parallel | Session 4 — Memory Hierarchies and Processing-in-Memory Chair: To be announced |
Session 5 — Memory, Communication, and System-Level Data Movement Chair: To be announced |
Session 6 — Accelerating AI Across GPUs and Mobile Systems Chair: To be announced |
| Talk 1 — #34: DREAM: In-DRAM Bit-Serial PIM with Data Reuse and Efficient Mapping Aman Arora; Jeeho Ryoo; Jiajun Hu; Lizy K. John; Siddhartha Raman Sundara Raman; Siyuan Ma |
Talk 1 — #318: Hermes: Accelerating Page Migration and HPC Data Transfers with NoC-attached Engines Adrià Armejach; Francesco Sgherzi; Ivan Vega; Jordi Fornt; Juan Miguel de Haro Ruiz; Marco Siracusa; Miquel Moreto; Pouya Esmaili Dokht |
Talk 1 — #195: Efficient Scheduling Algorithm for Large-scale Models on Heterogeneous Mobile Systems Jinyoung Kim; Minseong Kim; Yongjun Park; Yongjun Yongjun |
||
| Talk 2 — #73: SPARQ: Skew-Aware PIM Accelerator for Relational Join and Select Queries Sabiha Tajdari; Anastasia Ailamaki; Sandhya Dwarkadas |
Talk 2 — #91: Cordelia: A Huffmanized Merkle Tree for Secure Memory Galy Sela; Iris Bahar; Maurice Herlihy; Samuel Thomas; Tali Moreshet |
Talk 2 — #253: PALRAC: Parallel Linear Recurrence Accelerator for Tree-based Speculative Decoding Hyuk-Jae Lee; Sangheon Lee; Xuan Truong Nguyen |
||
| Talk 3 — #489: Focus on What Matters: DRAM-CXL Hybrid Memory Management with PRISM Daniel Mosse; Fatemeh Golshan |
Talk 3 — #143: A Coordinated Approach to Transactional Data Structures Ahmed Hassan; Michael Spear; Yaodong Sheng |
Talk 3 — #26: Split-Posit Systolic Array: A Resource-Efficient Hardware Accelerator for High-Performance AI Workloads Arun M; Madhav Rao; Sneha Dandekar; Vaishnavi Sharma |
||
| Talk 4 — #548: ElaCache: Fine-Grain Dynamic Partitioning of LLCs and Coherence Directories in Multiprocessors Adam Morrison; Dingyuan Cao; Josep Torrellas; Neil Zhao |
Talk 4 — #21: Proba: A High-Performance, Low-Traffic Probabilistic Spatial Memory Streaming Prefetcher Jacky Wong; Sam Ainsworth; Yinting Huang |
Talk 4 — #107: HiNa-MoE: High-Performance, Non-Intrusive MoE Inference on CPUs with Matrix Engines Dezun Dong; Enda Yu; Jianbin Fang; Junwen Zhang; Weiling Yang; Zhe Bai |
||
| 3:00–3:30 PM | Coffee Break | Coffee Break | ||
| 3:30–4:15 PM | Technical Sessions 7–8 — Parallel | Session 7 — Compilation for Specialized Computing Chair: To be announced |
Session 8 — Efficient ML Serving and Heterogeneous Scheduling Chair: To be announced |
ACM SRC Poster Finalists Chair: Jay Lofstead, Sandia National Laboratories Presenters will be announced on Tuesday, October 20, following the PACT Poster Session and Poster Reception. Finalists for the ACM Student Research Competition (SRC) will be selected based on the quality of their poster presentations and discussions during the Poster Session. Selected finalists will be invited to present their work in this session. |
| Talk 1 — #277: Reducing Address Arithmetic Overheads: New Compiler Techniques for Programmable Dataflow-based AI Accelerators Alberto Mannari; Alex Gatea; Bardia Mahjour; Chris Bowler; Masoud Ataei Jaliseh; Nicole Khoun; Prasanth Chatarasi; Shubham Jain; Swagath Venkataramani; Viji Srinivasan; Wei Wang |
Talk 1 — #519: HeteroSched: Co-Optimizing Scheduling and Parallelization for Deep Learning Workloads for Heterogeneous GPU Clusters Amirali Mirian; Bahram Afsharmanesh; Gagan Agrawal; Md Musfiqur Rahman Sanim |
|||
| Talk 2 — #58: Q-TranSim: Batch Quantum Circuit Simulation using Tensor Transpilation Hengrui Chen; Shui Jiang; Tsung-Wei Huang; Tsung-Yi Ho |
Talk 2 — #633: EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models Gagan Agrawal; Hangyu Zheng; Kunxiong Zhu; Miao Yin; Minghai Qin; Wei Niu; Zhihao Shu |
|||
| 4:15–5:00 PM | Industry Keynote — Samantika Sury, HPE | Plenary — Discovery Room. Chair: Antonino Tumeo, PNNL | ||
| 6:30–10:00 PM | PACT 2026 Banquet — Skydeck Chicago PACT Awards presented during the banquet |
|||
| Time | Program | Discovery Room | Classroom A | Classroom C |
|---|---|---|---|---|
| 8:00–9:00 AM | Registration | Registration / Check-in | ||
| 9:00–10:00 AM | Keynote — Mary Hall | Plenary — Discovery Room. Chair: Tanu Malik, University of Missouri, Columbia | ||
| 10:00–10:30 AM | Coffee Break | Coffee Break | ||
| 10:30 AM–12:00 PM | Technical Sessions 9–11 — Parallel | Session 9 — GPU Execution and AI Workload Optimization Chair: To be announced |
Session 10 — GPU Memory Systems and Data-Intensive Acceleration Chair: To be announced |
Session 11 — Performance Optimization, Scheduling, and Communication Chair: To be announced |
| Talk 1 — #141: Improving Data Reuse across Blocks for Efficient Block-sparse Transformers on GPUs Lihan Hu; Peng Jiang; Sun Xian-He; Xian-He Sun; Xiaoyang Lu |
Talk 1 — #39: SLGS: A Structure-Aware Scanline Renderer for Efficient 3D Gaussian Splatting Dongho Ha; Hyunwuk Lee; Mingu Jung; Seunghyun Lee; Sungbin Kim; Sungwoo Kim; Won Woo Ro; Yingyan (Celine) Lin |
Talk 1 — #97: Versatile Power Benchmark Generation for Modeling Emerging GPU Workloads and their Power Events Allison Seigler; Lizy K. John; Zhixing Jiang |
||
| Talk 2 — #75: Fast Cross-Operator Optimization of Attention Dataflow on Spatial Accelerators Bo Yuan; Hailiang Hu; Haodong Chang; Jiang Hu; Rongjian Liang; Yu Gong; Zhenrui Wang; Zhexiang Tang |
Talk 2 — #191: GZswap: Hardware-Managed Compressed Swap for GPU Memory Oversubscription Boyeol Choi; Jungrae Kim; Sanghyun Hong; Sanghyun Park; Seokin Hong |
Talk 2 — #475: HSF: A Hierarchical Scheduling Framework for Application-Tailored Scheduling on Tasking Runtimes Antoni Navarro Muñoz; David Álvarez; Vicenç Beltran; Vincent A. Arcila Larrea |
||
| Talk 3 — #327: KerColle: Unlocking Fine-Grained GPU Concurrency in Vision-Language-Action Models Anna Li; Christina Giannoula; Nandita Vijaykumar |
Talk 3 — #13: V-PWC: Accelerating Page Table Walk for Multi-Chip-Module GPUs Mingxin Tang; Qiong Li; Sen Yang; Wei Chen; Xia Zhao; Xinjin Gao |
Talk 3 — #146: Scalable Asynchronous Aggregation for Many-to-Many Transfers Naveen Namashivayam Ravichandrasekaran; Pen-Chung Yew |
||
| Talk 4 — #405: Mask-Aware Execution for Efficient JEPA Training Amirali Mirian; Bahram Afsharmanesh; Gagan Agrawal; Md Musfiqur Rahman Sanim; Wei Niu; Zhihao Shu |
Talk 4 — #238: CAPO: Enhancing CXL-based Memory Expanders with Adaptive Prefetching Anwen Huang; Chenglong Li; Jingyan Song; Na Nie; Qiong Li; Yongji Liu; Zhijie Liu |
Talk 4 — #190: VaBO: Delivering Reliable Autotuning Performance in Parallel Applications under Unreliable Conditions Hengrui Luo; Saeyeon Kim; Tirthak Patel; Vedica Rao; Younghyun Cho |
||
| 12:00–12:10 PM | Transition to Discovery Room — Please move to Discovery Room for the conference closing | |||
| 12:10–12:20 PM | PACT 2026 Closing Remarks | Discovery Room | ||
Program notes: All times are Central Time (CT). Session chairs and the poster-reception location will be updated when confirmed. Schedule subject to change.
Tuesday, October 20, 2026
Andrew A Chien
Univ of Chicago and Argonne National Lab
Chicago UpDown Computing, Inc.
Supercomputers and AI compute are ill-suited for sparse and data-intensive computations because they are optimized for maximum dense matrix "FLOPS". We have designed the UpDown System – optimized for graph computing, streaming data ingestion and transformation as well as high-level programming. The result outperforms conventional CPU/GPU-based systems by 10-100x on an ISO-power basis.
UpDown's radical micro-architecture unleashes fine-grained parallelism: 1-cycle thread creation and management, 1-cycle messages. This enables efficient computation on 10-instruction thread invocations. Further, software-controlled split-transaction DRAM access unlocks the power of HBM's massive memory bandwidth. For irregular applications, UpDown datapath efficiency is 10x greater. UpDown performance on skewed-graph computations exceeds multicore CPU's (>100x) and GPU's (20-60x) in single-node configuration. Updown performance scales to 1,000 and 10,000-fold speedup on BFS, Pagerank, Triangle Count, K-truss and more.
UpDown data ingestion exceeds 5 billion/records/s/node (1000x CPU-based databases), reaching 100 trillion records/s. It enables a new class of streaming analytics and complex workflows. UpDown enables high level programming with a global address space, and a flexible map-reduce framework (KVMSR) coupled with an event-driven language (UDWeave). This enables easy vertex, edge-centric programming, and fits well for other data-parallel models such as relational/graphDB, sparse matrixes, and more.
UpDown was created under funding from IARPA's AGILE program, and is being commercialized by Chicago UpDown Computing, Inc. (www.chupdown.com).
Andrew A Chien is the William Eckhardt Distinguished Service Professor of Computer Science at the University of Chicago and Senior Scientist at Argonne National Laboratories. Chien led the IARPA funded "UpDown System Project", designing breakthrough scalable graph analytics systems and is now Founder and President of Chicago UpDown Computing, Inc. (www.chupdown.com). He has led the Zero-carbon Cloud project since 2015, and is known for his research on datacenters, renewable energy and sustainability, cloud resource management and software, and large-scale system architecture. Chien has received numerous recognitions for research. Dr. Chien currently serves on the NSF CISE Advisory Committee and DARPA ISAT. He is a Fellow of the ACM, IEEE, and AAAS. He served as EiC of Communications of the ACM, 2017-2022, and Vice President of Research at Intel Corporation from 2005-2010. He served as SAIC Chair Professor of University of California, San Diego (1998-2005) and as faculty at the University of Illinois (1990-98). He received BS, MS, and PhD degrees from the Massachusetts Institute of Technology.
Who Gets to Do Computing Research in 2036?
Panelists: Eun Jung (EJ) Park — Qualcomm Innovation Center
Valerio Pascucci — University of Utah
Hariharan Devarajan — Lawrence Livermore National Laboratory
Tanu Malik — University of Missouri, Columbia
Lawrence Rauchwerger — UIUC
Moderator: Michela Taufer — University of Tennessee, Knoxville
Computing research is becoming increasingly expensive, complex, and concentrated. Emerging areas such as artificial intelligence and quantum computing often require specialized infrastructure, large datasets, substantial funding, technical staff, and extensive institutional capacity. Ambitious national initiatives such as the Genesis Mission further demonstrate the growing importance of coordinated research across academia, industry, national laboratories, and government. Yet the resources needed to participate in such efforts remain unevenly distributed.
Looking toward 2036, where will computing research take place, and who will be able to participate? Will the most consequential research become concentrated within a small number of well-resourced universities, companies, and national laboratories? What roles will smaller academic institutions, emerging companies, and individual researchers play?
This panel will bring together perspectives from academia, industry, and national laboratories to examine what counts as computing research, whether funding and infrastructure have become proxies for research excellence, and how institutional resources shape who can contribute. Panelists will discuss how cross-sector partnerships and national initiatives such as Genesis can broaden participation while preserving pathways for small teams, foundational and exploratory work, undergraduate-driven research, and institutions serving diverse students and regions.
Ultimately, the panel asks: How can we ensure that computing research in 2036 is shaped by the breadth and quality of its ideas—not only by where researchers work or the resources available to them?
Wednesday, October 21, 2026
Josep Torrellas
Thomas M. Siebel Chair in Computer Science
Director, SRC JUMP 2.0 ACE Center for Evolvable Computing
University of Illinois, Urbana-Champaign
iacoma.cs.uiuc.edu/josep/torrellas.html
Given current energy-consumption trends, there is ample consensus that we will have to move much of the computation to hardware accelerators. This is because accelerators are the most energy-efficient platforms. However, from the evidence of past efforts in this direction, architecting an accelerator-centric computing environment looks very challenging. It is unclear what architectural designs and software advances will really enable this new paradigm. In this talk, I will outline our vision of the hardware and software needed for a successful accelerator-centric computing environment, and some of the efforts that we are doing in this direction.
Josep Torrellas is the Thomas M. Siebel Chair in Computer Science at the University of Illinois, Urbana-Champaign (UIUC). He is the Director of the ACE Center for Evolvable Computing (an SRC/DARPA JUMP 2.0 Center), past Co-Leader of an Intel Strategic Research Alliance (ISRA) on Computer Security, and past Director of the Illinois-Intel Parallelism Center (I2PC). His research interests are multiprocessor computer architectures and parallel computing. Some of his contributions include thread-level speculation (TLS) architectures, the Bulk Multiprocessor concept, deterministic record and replay mechanisms, process variation mitigation techniques, and hardware defenses against speculative execution attacks. In addition, he has contributed to several experimental multiprocessor designs such as IBM's PERCS Multiprocessor, Intel's Runnemede Extreme-Scale Multiprocessor, Illinois Cedar, and Stanford DASH.
Torrellas has received the IEEE Computer Society (CS) Harry H. Goode Memorial Award, the UIUC Daniel C. Drucker Eminent Faculty Award, the UIUC Campus Award for Excellence in Graduate Student Mentoring, the IEEE CS Edward J. McCluskey Technical Achievement Award, and was a Willett Faculty Scholar at UIUC. He is an IEEE CS Golden Core Member, and a Fellow of IEEE, ACM, and AAAS. He was the Chair of the IEEE Technical Committee on Computer Architecture (TCCA). He has served in the Board of Directors of the Computing Research Association (CRA) and has been a Council Member of CRA's Computing Community Consortium (CCC). He was a member of the U.S. National Academies Board on Army Research and Development. He serves in the International Roadmap for Devices and Systems (IRDS). Torrellas has graduated 53 PhDs. He received a PhD from Stanford University.
Wednesday, October 21, 2026
Samantika Sury
Fellow and Chief Hardware Architect
HPE - HPC and AI Infrastructure Solutions
The end of Moore's Law scaling and the rapid rise of AI are driving a fundamental shift in computer architecture, accelerating the adoption of purpose-built technologies across compute, memory, networking, and storage. At the same time, scientific computing is moving beyond isolated applications toward tightly integrated workflows that combine simulation, data analytics, learning, inference, and increasingly agentic forms of execution. Together, these changes are reshaping how systems are designed and where performance bottlenecks emerge.
Next-generation AI-HPC systems will depend on tightly coupled and customizable architectures that bring specialized resources together around the needs of complete workflows rather than individual applications. This keynote explores architectural directions including workflow-centric optimization, flexible scale-up and scale-out fabrics, and macroheterogeneity. It will examine the trends driving these changes, the challenges they introduce, and the opportunities they create as the community moves toward more integrated, adaptable, and workload-aware AI-HPC systems.
Samantika Sury serves as an HPE Fellow, Vice President, and Chief Hardware Architect for HPC and AI Infrastructure Solutions. She leads the Future Technologies team, which focuses on advancing hardware and software system innovations. Samantika has previously held prominent roles at Samsung, where she served as Vice President and Chief Hardware Architect for HPC. She has also worked at Intel® as a Senior Principal Engineer, driving silicon and system architecture innovations into marketable products, served as the Chief Architect of Intel's HPC-Custom Silicon Program and was the Principal Investigator and Lead Architect for the DOE PathForward Program. Samantika holds 31 U.S. and international patents, has published more than 20 peer-reviewed papers, and has delivered numerous invited talks. She was recognized in HPCWire People to Watch 2026. Samantika earned her Ph.D. in Computer Science from the Georgia Institute of Technology.
Thursday, October 22, 2026
Mary Hall
Professor, Kahlert School of Computing, University of Utah
Data movement is the dominant execution and energy cost across the application workloads in data centers and supercomputers. Programming at the tile level has become a popular strategy for optimizing data movement for both deep learning and general structured grids, using Triton, cuTile, bricks, and fine-grained data blocks. Expressing hierarchical data and thread layouts, mostly designed with matrix processors in mind, facilitates automatic code generation that further raises the level of abstraction in such code. In this talk, we will describe prior work on BrickLib supporting fine-grained data blocks and active research on LEGO for hierarchical data and thread layout. We will connect these concepts with emerging hardware features and future demands on programming systems to reduce data movement.
Mary Hall is a Professor and former Director of the Kahlert School of Computing at University of Utah. Her research focuses on high-performance computing, compiler optimizations and code generation for novel and emerging hardware, and performance tuning. She has served on the Board of Directors of the Computing Research Association since 2015, and she is currently its Vice Chair. She is an ACM Distinguished Scientist and an IEEE Fellow.
| Title | Authors |
|---|---|
| PIM-SHERPA: Software Method for On-device LLM Inference by Resolving PIM Memory Attribute and Layout Inconsistencies | Sunjung Lee (Samsung Advanced Institute of Technology); Sanghoon Cha (Samsung Advanced Institute of Technology); Hyeonsu Kim (Samsung Advanced Institute of Technology); Seungwoo Seo (Samsung Advanced Institute of Technology); Yuhwan Ro (Samsung Advanced Institute of Technology); Sukhan Lee (Samsung Electronics); Byeongho Kim (Samsung Electronics); Yongjun Park (Yonsei University); Kyomin Sohn (Samsung Electronics); Seungwon Lee (Samsung Advanced Institute of Technology); Jaehoon Yu (Samsung Advanced Institute of Technology) |
| V-PWC: Accelerating Page Table Walk for Multi-Chip-Module GPUs | Sen Yang (College of Computer Science and Technology, National University of Defense Technology); Wei Chen (College of Computer Science and Technology, National University of Defense Technology); Xinjin Gao (College of Computer Science and Technology, National University of Defense Technology); Mingxin Tang (College of Computer Science and Technology, National University of Defense Technology); Qiong Li (Defense Innovation Institute); Xia Zhao (Defense Innovation Institute) |
| Proba: A High-Performance, Low-Traffic Probabilistic Spatial Memory Streaming Prefetcher | Yinting Huang (Huawei Newton Research Centre); Jacky Wong (Imperial College London); Sam Ainsworth (University of Edinburgh) |
| Split-Posit Systolic Array: A Resource-Efficient Hardware Accelerator for High-Performance AI Workloads | Sneha Dandekar (IIIT Bangalore); Arun M (IIIT Bangalore); Vaishnavi Sharma (IIIT Bangalore); Madhav Rao (IIIT Bangalore) |
| DREAM: In-DRAM Bit-Serial PIM with Data Reuse and Efficient Mapping | Siyuan Ma (University of Texas at Austin); Jiajun HU (Arizona State University); Jeeho Ryoo (Fairleigh Dickinson University); Siddhartha Raman Sundara Raman (The University of Texas at Austin); Aman Arora (Arizona State University); Lizy K. John (University of Texas at Austin) |
| SLGS: A Structure-Aware Scanline Renderer for Efficient 3D Gaussian Splatting | seunghyun lee (Yonsei University); Mingu Jung (Yonsei University); Sungwoo Kim (Yonsei University); Sungbin Kim (Yonsei University); Dongho Ha (Meta Platforms); Hyunwuk Lee (Unaffiliated); Yingyan (Celine) Lin (Georgia Institute of Technology); Won Woo Ro (Yonsei University) |
| AIEHalide: Compiling Halide to Spatial NPU Dataflow with Constrained Autoscheduling | Abnikant singh (IIIT Hyderabad); Abnikant Singh (AMD, Inc); Suresh Purini (IIIT Hyderabad) |
| Q-TranSim: Batch Quantum Circuit Simulation using Tensor Transpilation | Shui Jiang (The Chinese University of Hong Kong); Hengrui Chen (Zhejiang University); Tsung-Yi Ho (The Chinese University of Hong Kong); Tsung-Wei huang (UW Madison) |
| SPARQ: Skew-Aware PIM Accelerator for Relational Join and Select Queries | Sabiha Tajdari (University of Virginia); Anastasia Ailamaki (École Polytechnique Fédérale de Lausanne); Sandhya Dwarkadas (University of Virginia) |
| Fast Cross-Operator Optimization of Attention Dataflow on Spatial Accelerators | Haodong Chang (Texas A&M University); Hailiang Hu (AMD Inc); Zhenrui Wang (Texas A&M University); Yu Gong (Amazon Web Services); Rongjian Liang (Nvidia); Zhexiang Tang (Rutgers University); Bo Yuan (Rutgers University); Jiang Hu (Texas A&M University) |
| PDiMC: Achieving High-Throughput LLM Inference and Resolving Memory Concurrency via an Efficient PIM Subsystem | Byeongho Kim (Samsung Electronics); Kyomin Sohn (Samsung Electronics); Sukhan Lee (Samsung Electronics); Hweesoo Kim (Samsung Electronics); Wontak Han (Samsung Electronics); Yoonah Paik (Samsung Electronics); Jaewan Choi (Samsung Electronics) |
| Cordelia: A Huffmanized Merkle Tree for Secure Memory | Samuel Thomas (Pomona College); Galy Sela (EPFL); Tali Moreshet (Boston University); Maurice Herlihy (Brown University); Iris Bahar (Colorado School of Mines) |
| Versatile Power Benchmark Generation for Modeling Emerging GPU Workloads and their Power Events | Allison Seigler (University of Texas at Austin); Zhixing Jiang (University of Texas at Austin); Lizy K. John (University of Texas at Austin) |
| HiNa-MoE: High-Performance, Non-Intrusive MoE Inference on CPUs with Matrix Engines | Weiling Yang (National University of Defense Technology); Junwen Zhang (National University of Defense Technology); Dezun Dong (National University of Defense Technology); Jianbin Fang (National University of Defense Technology); Enda Yu (National University of Defense Technology); Zhe Bai (National University of Defense Technology); Xiaopeng Deng (National University of Defense Technology) |
| A Hybrid Processing-in-Memory Architecture for Long Sequence LLM Inference with KV Cache Filtering | Soojin Hwang (ETRI); Sanghyeon Lee (KAIST); Juhyun Lee (KAIST); Jaehyuk Huh (KAIST) |
| FIFO Initialization: Efficient Support for Serial Loops on Spatial Elastic CGRAs | Eric Xu (University of Toronto); Tarek S. Abdelrahman (University of Toronto) |
| ESA: Improving GPU Utilization with Elastic Isolation for ML Inference Services | Taeklim Kim (The University of Texas at Austin); Saurabh Agarwal (The University of Texas at Austin); Rachata Ausavarungnirun (MangoBoost Inc.); Jayneel Gandhi (Meta); Christopher J. Rossbach (UT Austin and Microsoft) |
| Improving Data Reuse across Blocks for Efficient Block-sparse Transformers on GPUs | Lihan Hu (The University of Iowa); Xiaoyang Lu (Illinois Institute of Technology); Xian-He Sun (Illinois Institute of Technology); Peng Jiang (The University of Iowa) |
| A Coordinated Approach to Transactional Data Structures | Yaodong Sheng (Lehigh University); Leoul Demissie (Lehigh University); Ahmed Hassan (Lehigh University); Michael Spear (Lehigh University) |
| Scalable Asynchronous Aggregation for Many-to-Many Transfers | Naveen Namashivayam Ravichandrasekaran (University of Minnesota); Nathan Wichmann (Hewlett Packard Enterprise); Pen-Chung Yew (University of Minnesota) |
| Stream Swizzling: Extending HLS for Efficient Non-Affine Stream Permutations | Chengyue Wang (UCLA); JONATHAN XUE (UCLA); LANCE GIANG (UCLA); SHINJU JU (UCLA); Yingquan Wu (MBZU AI Lab); Jason Cong (UCLA) |
| Splyce: SIMD Vectorization of Sparse Coiteration | Kabilan Mahathevan (Virginia Tech); Poorna Gunathilaka (Virginia Tech); Kirshanthan Sundararajah (Virginia Tech) |
| VaBO: Delivering Reliable Autotuning Performance in Parallel Applications under Unreliable Conditions | Vedica Rao (Santa Clara University); Saeyeon Kim (Santa Clara University); Hengrui Luo (Rice University); Tirthak Patel (Rice University); Younghyun Cho (Santa Clara University) |
| GZswap: Hardware-Managed Compressed Swap for GPU Memory Oversubscription | Boyeol Choi (Sungkyunkwan University); Sanghyun Park (FuriosaAI); Sanghyun Hong (Sungkyunkwan University); Seokin Hong (Sungkyunkwan University); Jungrae Kim (Sungkyunkwan University) |
| Efficient Scheduling Algorithm for Large-scale Models on Heterogeneous Mobile Systems | Jinyoung Kim (Yonsei University); Yongjun Kim (Samsung Electronics); Minseong Kim (Samsung Electronics); Yongjun Park (Yonsei University) |
| CAPO: Enhancing CXL-based Memory Expanders with Adaptive Prefetching | Jingyan Song (Academy of Military Sciences); Chenglong Li (Academy of Military Sciences); Anwen Huang (Academy of Military Sciences); Qiong Li (Academy of Military Sciences); Yongji Liu (Academy of Military Sciences); Zhijie Liu (Academy of Military Sciences); Na Nie (Academy of Military Sciences) |
| ST-Flow: A Hardware Compiler for Automating Spatial-Temporal Dataflow Acceleration | Suhail Basalama (UCLA); Stéphane Pouget (University of California, Los Angeles); Jason Cong (UCLA) |
| PALRAC: Parallel Linear Recurrence Accelerator for Tree-based Speculative Decoding | Sangheon Lee (Seoul National University); Hyuk-Jae Lee (Seoul National University); Xuan Truong Nguyen (Seoul National University) |
| KERYX: A CUDA/HIP Framework for Adaptive Runtime Compilation in Heterogeneous Systems | MARC GONZALEZ TALLADA (Universitat Politecnica de Catalunya); PEDRO VALERO (ORNL); KEITA TERANISHI; JEFF VETTER (ORNL) |
| Reducing Address Arithmetic Overheads: New Compiler Techniques for Programmable Dataflow-based AI Accelerators | Prasanth Chatarasi (IBM Research); Nicole Khoun (IBM); Wei Wang (IBM); Chris Bowler (IBM); Alex Gatea (IBM); Shubham Jain (IBM Research); Masoud Ataei Jaliseh (IBM); Alberto Mannari (IBM); Bardia Mahjour (IBM); Viji Srinivasan (IBM (Research)); Swagath Venkataramani (IBM Research) |
| Hoppolyta: Polyhedral Kernel Generation Meets Hopper Architecture | Aravind Acharya (NVIDIA); Somashekaracharya G Bhaskaracharya (NVIDIA); Evghenii Gaburov (NVIDIA); Bin Fan (NVIDIA); Alexander Collins (NVIDIA); Bastian Hagedorn (NVIDIA); Vinod Grover (NVIDIA) |
| Hermes: Accelerating Page Migration and HPC Data Transfers with NoC-attached engines | Francesco Sgherzi (SiPearl, Barcelona Supercomputing Center); Juan Miguel de Haro Ruiz (Barcelona Supercomputing Center); Jordi Fornt (Barcelona Supercomputing Center); Pouya Esmaili Dokht (Barcelona Supercomputing Center); Marco Siracusa (Barcelona Supercomputing Center); Ivan Fernandez (Barcelona Supercomputing Center); Adrià Armejach (UPC/BSC); Miquel Moreto (UPC/BSC) |
| KerColle: Unlocking Fine-Grained GPU Concurrency in Vision-Language-Action Models | Anna Li (University of Toronto); Christina Giannoula (Max Planck Institute for Software Systems (MPI-SWS)); Nandita Vijaykumar (University of Toronto) |
| SALT: Symbolic Analysis of Loop Tiling | Yanghui Wu (University of Rochester); Yifan Zhu (University of Rochester); Yekai Pan (University of Rochester); Chen Ding (University of Rochester) |
| Partial Instruction Execution on Long SIMD Architectures | Francesc Martinez (Barcelona Supercomputing Center); Adrià Armejach (UPC/BSC); Marc Casas (Barcelona Supercomputing Center) |
| ASTRA-MoE: GPU-Augmenting In-Storage Acceleration for Long-Context Mixture-of-Expert Inference | Hyeonggyu Jeong (Hanyang University); Inyoung Song (Hanyang University); Jinwoo Jeong (Hanyang University); Kyungmo Koo (Hanyang University); Byungmin Ahn (Samsung Electronics); Dong-Min Shin (Samsung Electronics); Yong Ho Song (Samsung Electronics); Jungwook Choi (Hanyang University) |
| Mask-Aware Execution for Efficient JEPA Training | Md Musfiqur Rahman Sanim (University of Georgia); Zhihao Shu (University of Georgia); Bahram Afsharmanesh (University of Georgia); Amirali Mirian (University of Georgia); Wei Niu (University of Georgia); Gagan Agrawal (University of Georgia) |
| Practical Correctness and Equivalence Checking for MLIR | Emily Tucker (Georgia Institute of Technology); Louis-Noel Pouchet (Colorado State University); Erika Hunhoff (AMD); Stephen Neuendorffer (AMD); Erwei Wang (AMD) |
| HSF: A Hierarchical Scheduling Framework for Application-Tailored Scheduling on Tasking Runtimes | Vincent A. Arcila Larrea (Barcelona Supercomputing Center); David Álvarez Robert (Barcelona Supercomputing Center); Antoni Navarro Muñoz (Barcelona Supercomputing Center); Vicenç Beltran (Barcelona Supercomputing Center) |
| Focus on What Matters: DRAM-CXL Hybrid Memory Management with PRISM | Fatemeh Golshan (University of Pittsburgh); Daniel Mosse (University of Pittsburgh) |
| HeteroSched: Co-Optimizing Scheduling and Parallelization for Deep Learning Workloads for Heterogeneous GPU Clusters | Bahram Afsharmanesh (University of Georgia); Md Musfiqur Rahman Sanim (University of Georgia); Amirali Mirian (University of Georgia); Gagan Agrawal (University of Georgia) |
| ElaCache: Fine-Grain Dynamic Partitioning of LLCs and Coherence Directories in Multiprocessors | Dingyuan Cao (University of Illinois at Urbana-Champaign); Neil Zhao (UT Austin & NVIDIA); Adam Morrison (Tel Aviv University); Josep Torrellas (Univ. of Illinois Urbana-Champaign) |
| EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models | Kunxiong Zhu (University of Georgia); Zhihao Shu (University of Georgia); Hangyu Zheng (University of Georgia); Minghai Qin (Western Digital Research); Miao Yin (University of Texas at Arlington); Gagan Agrawal (University of Georgia); Wei Niu (University of Georgia) |
| Title | Authors (Affiliations) |
|---|---|
| SAIL: SRAM-Accelerated LLM Inference System with Lookup-Table-based GEMV | Jingyao Zhang (University of California, Riverside); Jaewoo Park (Ulsan National Institute of Science and Technology); Jongeun Lee (Ulsan National Institute of Science and Technology); Elaheh Sadredini (University of California, Riverside) |
| BenchCPU: Performance Distribution-Aware CPU Benchmarking over Open Configuration Spaces | Chenxi Wang (Institute of Computing Technology, Chinese Academy of Sciences); Yuchen Su (State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences); Lei Wang (Institute of Computing Technology, Chinese Academy of Sciences); Guoxin Kang (University of Chinese Academy of Sciences); Wanling Gao (Institute of Computing Technology, Chinese Academy of Sciences); Fan Zhang (The International Open Benchmark Council); Jianfeng Zhan (Institute of Computing Technology, Chinese Academy of Sciences) |
| Dynamic Scheduling of LLM Inference in Asymmetric Memory Systems for Energy-Performance Balance | Soojin Hwang (ETRI); Jungwoo Kim (Stanford University); Sanghyeon Lee (KAIST); Hongbeen Kim (KAIST); Jaehyuk Huh (KAIST) |
| ABS: Efficiently Serving Deep Learning Models with Offline CVMs | Bojian Zheng (University of Toronto, CentML, Vector Institute); Wenyang Wang (University of Toronto, CentML, Vector Institute); Hui Shi (Tencent); Jie Fu (Tencent); Xiaofeng Yang (Tencent); Yangyu Tao (Tencent); Peng Chen (Tencent); Jie Jiang (Tencent) |
| FineCAT: Instruction-Count-Driven Fine-Grained LLC Management for Co-Location Scenarios | Yanqi Kan (Institute of Computing Technology, Chinese Academy of Sciences); Lei Wang (Institute of Computing Technology, Chinese Academy of Sciences); Fanda Fan (University of Chinese Academy of Sciences); Yikang Yang (Institute of Computing Technology, Chinese Academy of Sciences); Wanling Gao (Institute of Computing Technology, Chinese Academy of Sciences); Chunjie Luo (Institute of Computing Technology, Chinese Academy of Sciences); Jianfeng Zhan (Institute of Computing Technology, Chinese Academy of Sciences) |
| HARMONY: A Framework for Cooperative Memory Scheduling in CXL Memory Systems | Yongho Lee (Sungkyunkwan University); Junbum Park (Sungkyunkwan University); Sungbin Jang (Sungkyunkwan University); Osang Kwon (Samsung Electronics); Minkyu Choi (Samsung Electronics); Seokin Hong (Sungkyunkwan University) |
| HERO: Local HBM Enhancements to Support Remote Memory Optimizations | Christin David Bose (Purdue University); Cesar Avalos (Purdue University); Yechen Liu (Purdue University); Timothy Rogers (Purdue University) |
| Automating Exploration and Code Generation for Pipelined Dataflows on AMD XDNA™ NPUs | Joren Dumoulin (KU Leuven); Arne Symons (KU Leuven); Erika Hunhoff (AMD); Andre Roesti (AMD); Andra Bisca (AMD); Gagandeep Singh (AMD); Kristof Denolf (AMD); Marian Verhelst (KU Leuven) |
| TPE: AI-MicroBMT: A Unified Benchmark Tool for Deployment-Aware Performance Characterization of Diverse AI Accelerators | Jonghyun Shin (Seoul National University); Dongmyong Shin (Seoul National University); Jeongnam Youn (Suwon Science College); Dae-Hwan Kim (Seoul National University); Soojung Ryu (Seoul National University); Xuan Truong Nguyen (Seoul National University); Hyuk-Jae Lee (Seoul National University) |
| Diffusion-Based Data Augmentation for Multi-Label Performance Modeling | Mohammad Ali (Texas State University); Apan Qasem (Texas State University) |
| UniFlow: A Spatial Transformer Accelerator with Unified GEMM and Non-GEMM Dataflow Architecture | Haocheng Xu (University of California, Irvine); Faraz Tahmasebi (University of California, Irvine); Rachid Karami (University of California, Irvine); Zhiheng Chen (University of California, Irvine); Hyoukjun Kwon (University of California, Irvine); Sitao Huang (University of California, Irvine) |
| NDS: Programmer-Free Offload of High-Performance Near Data Strands | Shreyas Singh (University of Utah); Pratyush Nandi (University of Utah); Lin Jia (Intel); Shankar Balachandran (University of Utah); Rajeev Balasubramonian (University of Utah) |
| Cooperative Wavefronts: WFA and GWFA on a 4096-PE MIMD Many-Core | Wenjie Geng (University of Michigan - Ann Arbor); Noah Kaplan (University of Michigan - Ann Arbor); Reetuparna Das (University of Michigan - Ann Arbor); Nathaniel Bleier (University of Michigan - Ann Arbor) |
| SCISSOR: Scalable I/O for Small Scattered Objects Runtime | Kevin Assogba (Rochester Institute of Technology); Nigel Tan (Los Alamos National Laboratory); M. Mustafa Rafique (Rochester Institute of Technology); Michela Taufer (University of Tennessee Knoxville); Bogdan Nicolae (Argonne National Laboratory) |
| From Resampling to Retrieval: A Unified Budgeted View of Training with Large Candidate Sets | Ziyang Jia (University of Missouri, Columbia); Tanu Malik (University of Missouri, Columbia) |
| Title | Authors (Affiliations) |
|---|---|
| RoadBlock: Rethinking GPU Tensor Core Microarchitecture for Emerging Microscaling Format Support | Nikhil Rout (University of California, Los Angeles, United States) |
| Diagnosing GPU Performance Regressions Across Compiler Versions with Strata | Befikir Bogale (University of Tennessee, United States) |
| Software-Controlled GB-Scale 3D-SRAM Residency for GPUs | Eric Dubberstein (Carnegie Mellon University, United States) |
| Causal Observability for Microarchitectural Analysis | Saber Ganjisaffar (University of California, Riverside, United States) |
| BIMBA: Best-Effort In-Network Merging with Bank-Side Accumulation for Transformer Accelerators | Jungwoo Park (Seoul National University of Science and Technology, South Korea) |
| SpeedSparse: Exploiting Dense Matrix Multiplication for Practical Sparse Neural Networks | Shreya Alladi (Computer Engineering Department, University of Murcia, Spain) |
| Model-Scale-Dependent Effects of Thread-Count Scaling on INT8 Dynamic Quantization | Nahla Nabil Skaik (Arab Open University - Bahrain, Bahrain) |
| ACIO: Always-Complete Isolation of Outliers with Parallelism-Amortized Hardware for Low-Bit LLM Quantization | Jihyeon Hwang (Seoul National University of Science and Technology, South Korea) |
| Joint Bit-Width and Fan-In Sensitivity Optimization for FHE-Aware Neural Network Acceleration | Murat Toprak (Istanbul Technical University, Turkey) |
| Roofline-Decomposed Agents for Sample-Efficient On-Device LLM Execution | Kaiyuan Zhang (Univeristy of Georgia, United States) |
| XBM: Hybrid Stack Composition for 3D Memory-on-GPU Inference | Atharva Raut (Carnegie Mellon University, United States) |