skalujto.ai

Menu

  • AI
  • LLM
  • open-source
  • Qwen3
  • agentic coding
  • benchmarks

Qwen3.6-35B-A3B FP8 at 256k context — how does it compare to other LLMs?

Artur Niklewicz
Qwen3.6-35B-A3B FP8 at 256k context — how does it compare to other LLMs?

On April 16, 2026, Alibaba's Qwen team released Qwen3.6-35B-A3B under the Apache 2.0 licence — making one of the most capable coding-focused open-source models freely available to run locally. We've spent the last two weeks benchmarking the FP8-quantised version on our NVIDIA DGX Spark against the leading proprietary models. Here is what we found, and when it is genuinely worth choosing a local model over a cloud API.

Architecture: Why MoE changes the maths

Qwen3.6-35B-A3B is a sparse Mixture-of-Experts model. The numbers in the name tell you everything: 35B total parameters, but only 3 billion activated per token during inference. For comparison, a dense 7B model activates all 7B parameters on every forward pass. The MoE architecture means you get near-35B quality at a fraction of the compute cost — the key reason this model punches well above its inference weight class.

FP8 quantisation (8-bit floating point) halves the memory footprint compared to BF16 with negligible quality loss on NVIDIA Ada and Hopper GPUs. The FP8 weights sit at approximately 35 GB; at runtime with activation overhead, the model uses around 40–42 GB of VRAM. On an H100 80 GB this leaves 30+ GB for KV cache — enough to actually exploit the 262,144-token context window.

The 262k context window — and why it matters for agentic coding

Most coding benchmarks test single-function generation in isolation. Agentic coding is a fundamentally different task: the model must read an entire repository, trace call graphs across files, understand test expectations, and make edits that are internally consistent. A 128k context fits perhaps 80–100k tokens of Python. A 262k context fits a complete medium-sized project — source, tests, CI config and error log — in a single prompt. Qwen3.6-35B-A3B's native context can be extended beyond 1 million tokens via YaRN scaling, which matters for large monorepos.

Benchmark results: 73.4% on SWE-bench Verified

SWE-bench Verified is the closest thing to a real-world software engineering test available: the model is given a GitHub issue and must produce a working patch. Qwen3.6-35B-A3B scores 73.4% — positioning it alongside Claude Sonnet 4 in agentic coding tasks, and ahead of models several times larger when measured at inference cost. Evaluations used an internal agent scaffold with bash and file-edit tools, temperature 1.0, top_p 0.95, and the full 200k context window.

Our internal results across three task categories (algorithmic problems, multi-file FastAPI + Next.js refactoring, SWE-bench Verified instances):

  • vs Claude Sonnet 4: roughly on par for isolated algorithmic tasks; slightly behind on complex multi-step tool-use chains
  • vs GPT-4o: competitive on Python/FastAPI backend work; weaker on TypeScript and frontend-heavy tasks
  • vs Gemini 1.5 Pro: consistently edges it on Python-heavy tasks; roughly equivalent on mixed full-stack work

Running it on DGX Spark: setup and real numbers

We serve Qwen3.6-35B-A3B-FP8 via vLLM with CUDA FP8 kernels. The recommended launch command from the Qwen team:

vllm serve Qwen/Qwen3.6-35B-A3B-FP8 --port 8000 --tensor-parallel-size 8 --max-model-len 262144 --reasoning-parser qwen3

Throughput at full 256k context: approximately 15–25 tokens/second on the DGX Spark. With Multi-Token Prediction (MTP) speculative decoding enabled, the community reports up to 80 tok/s on smaller context windows. For our use cases — code review pipelines, CI-integrated quality checks, and internal tooling — the latency profile is entirely acceptable. Crucially: all inference stays on-premise. No token leaves our network.

When to use it — and when not to

Local Qwen3.6-35B-A3B FP8 is the right choice when:

  • Code contains proprietary business logic that legally cannot leave your infrastructure
  • A client NDA prohibits transmission to third-party cloud APIs
  • You need version-pinned, reproducible inference with zero model drift between runs
  • GDPR or other regulations require data to remain within a defined jurisdiction

Cloud models (Claude, GPT-4o) remain the better choice when you need broad world knowledge combined with coding ability, when latency is critical, or when the code does not touch sensitive data. These are not competing approaches — they serve different risk and capability profiles and we use both, depending on the project.

Verdict

Qwen3.6-35B-A3B FP8 is the strongest open-source coding model we have run locally. The 73.4% SWE-bench score is not a lab curiosity — it translates into genuine multi-file refactoring capability at 262k context that was simply not available in a local, privacy-preserving form six months ago. If your team handles NDA-protected or regulated code, and you have access to an H100 or DGX Spark, this model earns its setup cost.

See also