Research | arcSYSu Lab

Our Research

Architecture-aware Compilation and Systems for
Scalable HPC/AI

We build compilation and cross-layer systems that help HPC and AI applications adapt to evolving architectures, dynamic resources, and increasing scale.

A unified systems question

How can HPC/AI applications make effective use of modern computing infrastructure?

Program Architecture
Task Resource
Application Platform

Transform programs · Orchestrate resources · Scale applications

01 Heterogeneous architectures

New instructions, memory systems, and accelerators

02 Dynamic workloads

Changing demands, states, and resource constraints

03 Scalable execution

Portable environments, parallel devices, and services

Research directions


Three connected layers of HPC/AI infrastructure

Our research spans program transformation, runtime resource management, and complete application execution. Together, these layers address the heterogeneity, dynamism, and scale of modern HPC/AI infrastructure.

01

Transform programs

Program Transformation & Architecture Adaptation

How can programs adapt to rapidly evolving computing architectures?

We develop compilation and program-transformation techniques across data organization, intermediate representations, instruction streams, and binary code. Our goal is to help scientific and AI workloads exploit emerging architectural capabilities while remaining efficient and portable across platforms.

Core capability Architectural adaptability · Code efficiency · Hw/Sw co-design
Computation mapping Data movement IR transformation Cross-ISA translation
coMulator MICRO 2026

Compilation-assisted cross-architecture translation and emulation

GoPTX DAC 2025

PTX-level instruction interleaving for fine-grained GPU kernel fusion

HSPref MICRO 2026

Architecture-aware data supply prefetching for ARM SME outer-product workloads

HStencil SC 2025

Mapping stencil computation to ARM SME outer-product engines

02

Orchestrate resources

Adaptive Runtime Resource Orchestration

How can computating resources respond to changing program behavior and workload demand?

We manage kernels, loading decisions, GPU sharing, and service-level spatiotemporal allocation. Our runtime mechanisms move resource provision from fixed configuration toward workload-driven, state-aware, and adaptive orchestration.

Core capability Resource efficiency · Runtime responsiveness · Service quality
Resource extension Kernel scheduling On-demand loading GPU sharing Space–time orchestration
Bullet ASPLOS 2026

Dynamic spatial-temporal orchestration for improving GPU utilization in LLM serving

HuntKTm TACO 2025

Hybrid scheduling and automatic management for efficient GPU kernel execution

PaSK DAC 2025

Proactive and selective kernel loading for mitigating inference cold starts

SMILE DAC 2024

Extending GPU shared memory with last-level cache capacity

03

Scale applications

Portable and Scalable HPC/AI Execution

How can complete applications deploy easily and execute efficiently at scale?

We coordinate software environments, computation, communication, parallel pipelines, and global traffic across the application lifecycle. These mechanisms enable complex HPC/AI applications to move from portable deployment to efficient parallel execution and scalable online service.

Core capability Environment portability · Parallel efficiency · Distributed scalability
Environment reconstruction Compute–communication fusion Pipeline reconfiguration Global flow control
coMtainer SC 2025

Compilation-assisted reconstruction of HPC container images for cross-platform portability

FusedRec AAAI 2026

Compute–communication fusion for distributed recommendation training

DynaPipe NeurIPS 2025

Dynamic layer redistribution for efficient pipeline-parallel LLM serving

gLLM SC 2025

Global pipeline balancing and token-flow control for distributed LLM serving

From systems research to shared capability

YatCC: validation, platformization, and impact

YatCC translates our broader systems expertise into reusable AI-native workspaces and intelligent services. It serves as a living testbed for evaluating systems ideas with real users and workloads, while turning research experience into shared capabilities for scientific research, education, and engineering practice.

Visit YatCC
30+ AI models
30B+ tokens served
700+ container instances

Explore arcSYSu

From Architecture to Infrastructure —
Turning Insight into Impact.