Gemma 4 E2B hits 20 tokens/sec on phone

// 54d agoBENCHMARK RESULT

Gemma 4 E2B hits 20 tokens/sec on phone

A Reddit user says they ran Google’s newly launched Gemma 4 E2B fully offline on a phone and measured 20.3 tokens/sec on GPU inference. It is a small but useful real-world datapoint for Gemma 4’s edge positioning, suggesting the model is not just theoretically mobile-friendly but can feel practical on-device.

// ANALYSIS

Strong signal for on-device AI, but still an anecdotal benchmark from one setup.

–The result supports Google’s claim that Gemma 4 E2B is designed for mobile-first, offline use.
–20.3 tok/s on a phone is a credible interactive speed, not just a lab curiosity.
–The post does not specify device model, quantization, runtime, prompt length, or thermal conditions, so the number is not broadly comparable yet.
–As a community datapoint, it matters more for feasibility than for leaderboard-style benchmarking.

// TAGS

gemmagemma-4e2bon-device-aimobile-aioffline-inferencebenchmarkgoogle-deepmind

DISCOVERED

54d ago

2026-04-03

PUBLISHED

54d ago

2026-04-03

RELEVANCE

8/ 10

AUTHOR

EthanJohnson01

// KEEP READING

More AI developer news from the feed

EXPLORE FULL FEED

UPDATE2h ago

Cursor adds dedicated subagents for skills

Cursor now allows developers to execute tool-heavy or research-intensive agent skills within dedicated subagents. This architectural shift isolates noisy background tasks, keeping the main chat context clean and focused.

UPDATE2h ago

YouTube moves AI labels to video player

YouTube is moving its AI content disclosures from video descriptions to more prominent placements beneath the player and on Shorts overlays. Starting in May, the platform will use internal signals to automatically label photorealistic AI content that creators fail to disclose.

OPEN SOURCE5h ago

Taste Skill kills AI "frontend slop"

Taste-Skill is an open-source framework that provides portable "agent skills" to enforce high-end design principles in AI-generated code. By injecting specific design directives and "anti-slop" rules, it enables LLMs to produce editorial-grade UIs that bypass generic, boilerplate-heavy AI templates.