Other

SF Technology · HAMi

Vendor-reportedLogistics & transportChinaIn productionEnterpriseLLM

Vendor-reported. The customer is named and the numbers are quoted from the source page, but the account comes from the vendor. No independent confirmation.

Reported by dynamia.ai (vendor self-report). Checked against the source page on 2026-09-06. 4 of 4 figures below appear on that page word for word.

The problem

Traditional GPU usage patterns led to underutilization (below 30%), resource waste, coarse scheduling, and difficulties adapting to heterogeneous devices.

What was deployed

Vendor
HAMi
Products
EffectiveGPU, HAMi
Technique
LLM
Build or buy
Bought and customised
Deployment
Custom build
Scale
65 services on 28 GPUs; 19 services on 6 test GPUs

What changed

Each row is quoted from the source. Figures we could not find on the page in those words are marked — they are kept, not deleted, so you can judge them.

37 cardsGPU cards saved for large model inference services

Deployed 65 services using 28 GPU cards, saving 37 cards

Quoted word for word from the source

13 cardsGPU cards saved for testing service cluster

Deployed 19 services using 6 test GPU cards, saving 13 cards

Quoted word for word from the source

0.5%Performance degradation after adding pooling layer

Minimum performance decrease of only 0.5% after adding pooling layer

Quoted word for word from the source

up to 200%Memory overcommitment ratio

Introduced dual-dimension overcommitment technology for memory and computing power (up to 200% memory overcommitment ratio)

Quoted word for word from the source

EnsuredReal-time performance for critical tasks

Ensured real-time performance for critical tasks through priority scheduling and resource overcommitment

Quoted word for word from the source

Through close collaboration with the HAMi open-source community and secondary innovation based on its framework, EffectiveGPU has helped us significantly improve GPU resource efficiency and reduce operational costs. This is an exemplary case of win-win cooperation between open-source collaboration and enterprise practice.

Difficulties and limits

Sources

Similar deployments

ANZ Bank · NVIDIA

0.82 — Gini coefficient for assessing risk

Independently reportedBanking & finance · Other · Classical ML · Pilot

Unilever · Google Cloud

15-35% increase — retailer sales

Independently reportedOther · Other · Classical ML · Scaled

BHP Billiton

zero — accidents due to driver drowsiness

Independently reportedOther · Other · Classical ML · In production · 2022

Cubo Ai · Google Cloud

more than 10X — user growth supported with same IT workforce

Vendor-reportedOther · Other · Computer vision · Scaled · 2019