Qwen3.5-24B REAP squeezes agentic coding into 16GB

// 83d agoPRODUCT LAUNCH

Qwen3.5-24B REAP squeezes agentic coding into 16GB

A LocalLLaMA contributor released a 32% expert-pruned GGUF variant of Qwen3.5-35B-A3B aimed at coding and agentic workflows on lower-VRAM hardware. The release includes quantized checkpoints, pruning/quantization scripts, and a reproducible Modal pipeline.

// ANALYSIS

This is a practical community optimization drop, not a new base model, but it materially lowers the barrier to running strong MoE coding models locally.

–The model trims experts from 256 to 175 while keeping ~3B active parameters per token, targeting better memory efficiency.
–The recommended IQ4_K_S GGUF is positioned for 16GB-class GPUs, which is the core value proposition here.
–The author shares full replication assets (REAP fork + Modal scripts), making this useful for other quantizers and pruning experiments.
–Calibration limits (1024 context and memory pressure during profiling) suggest further quality/perf gains are still possible.

// TAGS

qwen3.5local-llmggufmoequantizationagentic-coding

DISCOVERED

83d ago

2026-03-05

PUBLISHED

83d ago

2026-03-04

RELEVANCE

8/ 10

AUTHOR

tubuntu2

// KEEP READING

More AI developer news from the feed

EXPLORE FULL FEED

UPDATE1h ago

Cursor adds dedicated subagents for skills

Cursor now allows developers to execute tool-heavy or research-intensive agent skills within dedicated subagents. This architectural shift isolates noisy background tasks, keeping the main chat context clean and focused.

UPDATE1h ago

YouTube moves AI labels to video player

YouTube is moving its AI content disclosures from video descriptions to more prominent placements beneath the player and on Shorts overlays. Starting in May, the platform will use internal signals to automatically label photorealistic AI content that creators fail to disclose.

OPEN SOURCE5h ago

Taste Skill kills AI "frontend slop"

Taste-Skill is an open-source framework that provides portable "agent skills" to enforce high-end design principles in AI-generated code. By injecting specific design directives and "anti-slop" rules, it enables LLMs to produce editorial-grade UIs that bypass generic, boilerplate-heavy AI templates.