Blog

Insights, tutorials, and news from our team

HAMi's Place in the CNCF × PyTorch Stack

Community

HAMi Meetup Shanghai Deep Dive (1): Where Does HAMi Fit in the CNCF × PyTorch Cloud Native AI Foundation?

The model runs, yet the business never ships; one team waits for GPUs while another's sit idle. In his opening keynote at HAMi Meetup Shanghai, CNCF China Director Keith Chan laid out a layered view of a cloud native AI R&D foundation — with HAMi in the scheduling and resource management layer, handling heterogeneous compute sharing. Understanding that position tells you what HAMi solves, and which components it needs alongside it.

Community Panel: From User to Co-Builder

Community

HAMi Meetup Shanghai Deep Dive (7): From User to Co-Builder — Why Enterprises Need a Sustainable HAMi Community

Once open source enters production, enterprise attention extends beyond the code: will issues get answers, how do needs reach the public roadmap? The HAMi Meetup Shanghai panel covered real-world usage, contribution barriers, responsibility boundaries in AI-assisted development, and knowledge preservation — a second set of selection criteria beyond technology fit.

HAMi 2.10 — Isolation, Scheduling, Observability

Community

HAMi Meetup Shanghai Deep Dive (2): After GPU Sharing — Isolation, Scheduling, and Observability in HAMi 2.10

Slicing one GPU across multiple jobs is only the beginning of sharing. Dynamia co-founder and CTO and HAMi Maintainer Li Mengxuan walked through HAMi v2.10 progress, DRA ecosystem adaptation, and the v2.11 roadmap announced on stage (Remote GPU, CPU NUMA alignment). The through-line: after sharing, compute must be allocated on demand and constrained, observable, and diagnosable while it runs.

Iluvatar CoreX GPU Inference in Practice

Community

HAMi Meetup Shanghai Deep Dive (4): Turning GPU Performance into Inference Service Capability — Lessons from Iluvatar CoreX

The GPU benchmarks well, yet users still wait for the first token or feel stalls mid-generation. Ye Chenglin, Infra R&D Director at Iluvatar CoreX (天数智芯), shared a full engineering path for domestic-GPU inference at HAMi Meetup Shanghai: execution optimization, topology-aware scheduling, and inference orchestration — accepted against end-to-end SLOs, with HAMi covering the small-model GPU-sharing branch.

llm-d and HAMi: Dividing the Work

Community

HAMi Meetup Shanghai Deep Dive (3): Inference Control vs. GPU Resource Management — How llm-d and HAMi Can Work Together

You deployed multiple replicas of the same model and round-robined requests, yet responses didn't get faster. Red Hat Greater China CTO Zhang Jiaju presented llm-d's distributed inference plane at HAMi Meetup Shanghai: how prefix-aware routing, PD disaggregation, and cache offloading cooperate — and, in the Q&A, the collaboration direction with HAMi. Request routing and GPU slicing are two layers that must connect, yet cannot replace each other.

UCloud's AI Platform Engineering

Community

HAMi Meetup Shanghai Deep Dive (6): How UCloud Turns HAMi Sharing into Ready-to-Use Development Environments

The GPUs arrived; the algorithm team is still waiting for environments. Peng Peng, Senior R&D Engineer at UCloud, shared "Engineering Challenges and Practices of AI Computing Platforms" at HAMi Meetup Shanghai: unified instance specs, multi-cluster access, tool injection, and image acceleration shorten the environment delivery chain — with HAMi vGPU for private deployments, and separate kernel-level vGPU and QEMU passthrough paths for public cloud.

iFLYTEK's Volcano + HAMi-core Practice

Community

HAMi Meetup Shanghai Deep Dive (5): How iFLYTEK Manages Multi-Business GPU Sharing with Volcano + HAMi-core

Research jobs want GPUs now; online business needs guaranteed resources. Dong Jiang, Senior Architect at iFLYTEK, shared "Building a K8s Heterogeneous AI Compute Base with Volcano + HAMi-core" at HAMi Meetup Shanghai: queue management and orchestration layered with in-card resource control, letting different businesses keep their own execution paths while shared resources get clear usage rules.

CETC Cloud CNCF Case Study

Case Study

CETC Cloud Builds a Domestic GPU Sharing Foundation for a Portable Knowledge Base with HAMi

CETC Cloud needed to bring its intelligent knowledge base to project sites, running text generation, embedding, rerank, knowledge processing, and R&D debugging on a single set of portable devices with domestic GPUs. With Kubernetes and HAMi, per-device dev environment capacity grew from 2 to 30 concurrent Pods, and a single card freed 56 GB of memory and 80% of compute.

Subscribe to our blog

Get the latest technical articles, tutorials, and HAMi community updates.

Blog - Dynamia AI | GPU Virtualization & HAMi