All reel notes
Reel note

Qwen 3.8 27B: Opus-level intelligence that runs locally on a 4090

Aug 28, 20263 min read@littlebittech
Qwen 3.8 27B: Opus-level intelligence that runs locally on a 4090

Man, I love Chinese AI labs. Qwen 3.8 - a 27 billion parameter model - is now on the market, and in several benchmarks it has the intelligence of Opus 4.6, which was a 500B+ parameter model.

The benchmark shock

Typed out as a table, it is wild seeing a 27B dense model trade blows with the frontier:

BenchQwen 3.8 27BOpus 4.6 Max
SWE-bench Pro61.753.4
OSWorld84.372.7
QwenSWEBench79.063.8
AndroidWorld81.962.0
LiveCodeBench90.388.8

Same SWE benchmarks as Opus. Mind-blowing.

Capabilities

Local hardware

Here is where it gets fun: with 4-bit quantization, Qwen 3.8 27B runs on an RTX 4090. Roughly 13-17 GB VRAM, ~41 tokens/sec. On hardware you already own.

Run it on a single 4090 - free compute of your own
Run it on a single 4090 - free compute of your own

The new paradigm

Compare the two worlds:

The frontier-vs-home intelligence gap is closing. I will take that look. Thank you, Qwen.

Reels like this, every few days.

Follow for short bites on AI, tools, and everything I am shipping.

Follow on Instagram