DeepSeek and Huawei Launch TileLang: The Open-Source CUDA Alternative That Could Lower AI Costs
Back to blog
AI Automation 7 min 606 wordsSeptember 30, 2026

DeepSeek and Huawei Launch TileLang: The Open-Source CUDA Alternative That Could Lower AI Costs

DeepSeek today released an open-source toolkit built with Huawei that includes TileLang —the first serious rival to Nvidia's CUDA monopoly— plus libraries like DeepGEMM, FlashMLA and DeepEP optimized for Ascend 950 chips. For SMBs, this could mean meaningfully lower AI compute costs within the next 12–18 months.

⚡SEE LIVE DEMOS

On September 30, 2026, DeepSeek took one of its boldest steps yet by publicly releasing a complete AI chip programming toolkit developed in close collaboration with Huawei Technologies. The package includes TileLang —a high-level language that acts as a direct alternative to CUDA, Nvidia's de facto standard that has dominated AI development for over a decade— alongside specialized libraries: DeepGEMM for matrix multiplication, FlashMLA for LLM inference, TileKernel, DeepSelect, and DeepEP. Everything is open-source, free to download, and optimized for the Huawei Ascend 950 chips. This move marks an inflection point in the AI hardware ecosystem, with direct implications for the costs that businesses —including SMBs— pay when using AI cloud services.

On September 30, 2026, DeepSeek took one of its boldest steps yet by publicly re

What Did DeepSeek and Huawei Actually Release?

Today's toolkit is not an academic experiment: it is production software that DeepSeek already used internally to run its own models on Huawei hardware. TileLang provides a high-level abstraction layer similar to what Nvidia's CUDA offers, meaning a developer familiar with the Nvidia ecosystem can migrate their models to Huawei Ascend chips without learning an entirely alien system. DeepGEMM and FlashMLA are low-level optimizations for the most compute-intensive operations in LLMs (transformers), while DeepEP handles inter-chip communication in multi-GPU configurations. Additionally, DeepSeek and Huawei co-developed a 128-chip Ascend 950 'supernode' solution that lets them compete in performance with Nvidia H100/H200 GPU clusters for training and inference of large models. The toolkit generalizes that earlier effort: any company or cloud provider can now attempt the same migration with their own models.

Today's toolkit is not an academic experiment: it is production software that De
"

"When the software layer connecting models to hardware stops being a monopoly, AI compute prices drop. Not tomorrow, but within 18 months. SMBs planning their AI infrastructure today have the opportunity to choose from more providers at better prices."

Davarion Group & Labs

Real Impact for SMBs

  • 01Lower inference costs in the medium term: Cloud providers that adopt Huawei Ascend 950 chips will be able to offer cheaper GPU-hours, potentially reducing the cost of running proprietary AI models or using APIs by 30–50% within 12–18 months.
  • 02More cloud AI provider options: Current reliance on AWS, Azure, and Google (which predominantly run on Nvidia hardware) will decrease as new data centers —especially in Latin America and Asia— adopt the DeepSeek + Ascend 950 stack.
  • 03Open-source does not mean unstable: TileLang, DeepGEMM, and the rest already ran production models at DeepSeek. The battle-tested code significantly reduces early adoption risk.
  • 04Immediate recommended action: If your business hosts AI models or plans to run proprietary models on-premise, ask your cloud provider whether they plan to support Ascend 950; early movers will be able to negotiate better compute contracts.

To understand the magnitude of this move, it helps to appreciate what CUDA represents: since 2007, Nvidia's software ecosystem has been the only practical path for training and inferring AI models at scale. Frameworks like PyTorch and JAX were built on top of CUDA. This gave Nvidia a near-unbreachable competitive moat —not because of the hardware itself, but because the cost of rewriting software for other chips was prohibitive. TileLang attempts to solve exactly that problem: by offering an abstraction compatible with CUDA patterns, it dramatically reduces migration friction. If the open-source community adopts it (which depends on whether Ascend chips become accessible outside China), downward pressure on AI compute pricing will intensify significantly.

To understand the magnitude of this move, it helps to appreciate what CUDA repre

At Davarion Group & Labs, we help businesses in Houston TX and across Latin America design their AI infrastructure independently of any single vendor. If your business currently spends more than 15% of its tech budget on AI APIs or cloud compute, it's worth reviewing your architecture before the hardware market shifts again. Contact us at davarion.com for an AI cost optimization consultation.

At Davarion Group & Labs, we help businesses in Houston TX and across Latin Amer
#DeepSeek#Huawei Ascend#TileLang#CUDA alternative#AI infrastructure#SMBs

Davarion Group & Labs

WANT TO SEE THE AI IN ACTION?

Try an AI chatbot configured with your business name — live, no signup required.