
TileLang v0.1.15 adds Ascend 950
TileLang v0.1.15 adds a native Huawei Ascend 950 backend with code generation, automatic Cube/Vector scheduling, synchronization, and mixed SIMD/SIMT kernels. It also brings opt-in CUDA warp specialization, unified block-scaled GEMM, and a more expressive Python frontend.
This is a strategically important release: TileLang is evolving from a GPU-focused DSL into a credible cross-accelerator kernel platform while preserving low-level performance control.
- –Ascend 950 support covers compilation, scheduling, memory scopes, synchronization, PyTorch NPU integration, profiling, and production-oriented examples.
- –Automatic CUDA warp specialization assigns TMA loads, MMA computation, stores, and worker tasks to dedicated warp groups, reducing manual pipeline engineering.
- –Unified `T.gemm_blockscaled` semantics give FP8/FP4 workloads a more portable programming model across supported backends.
- –Python compile-time iteration, `enumerate`, `zip`, comprehensions, and generator expressions make complex kernel-generation code easier to express.
- –The release includes compatibility changes, so teams upgrading from v0.1.14 should review removed parameters and backend-specific imports.
DISCOVERED
2h ago
2026-10-01
PUBLISHED
1d ago
2026-09-30
RELEVANCE