---
type: "post"
slug: "thinking-machines-releases-inkling-small"
title: "Inkling-Small opens a 12B-active multimodal reasoning model"
description: "Thinking Machines released full weights for a 276B-parameter MoE with 12B active parameters and a one-million-token context window."
retrieval_nugget: "Thinking Machines released full weights for a 276B-parameter MoE with 12B active parameters and a one-million-token context window."
published_at: "2026-08-16"
updated_at: "2026-08-16"
record_date: "2026-08-16"
date_kind: "scheduled_at"
topics: ["open-models","multimodal","reasoning","agents"]
entities: ["thinkingmachines.ai"]
editorial_format: "field_note"
basket_id: "64af3bcb-1c2d-42a9-a664-91510a61d75a"
basket_revision: 1
source_urls: ["https://thinkingmachines.ai/news/inkling-small"]
visual_decision: "text_only"
recovery_incident: "NR-2026-08-15-HERMES-SITE-COPY"
schema_version: "newruntime-agent-readable-v0.2"
stable_id: "post:thinking-machines-releases-inkling-small"
status: "published"
visuals: []
editorial_provenance: {"schema_version":"newruntime-editorial-copy-v1","content_status":"source_grounded_final","final_copy_sha256":"sha256:1b8ef287a23d35c530d4a4adc43e3a8762a85c1421796b169908c1d6ecb21d9d","reviewed_at":"2026-08-15T20:30:00.000Z","source_evidence_count":1,"verified_claim_count":2}
routes: {"html":"https://newruntime.com/posts/thinking-machines-releases-inkling-small/","markdown":"https://newruntime.com/posts/thinking-machines-releases-inkling-small.md","json":"https://newruntime.com/posts/thinking-machines-releases-inkling-small.json"}
---

# Inkling-Small opens a 12B-active multimodal reasoning model

## Retrieval answer

Thinking Machines released full weights for a 276B-parameter MoE with 12B active parameters and a one-million-token context window.

Thinking Machines released the full weights of Inkling-Small, a Mixture-of-Experts model with 276 billion total parameters and 12 billion active parameters. It supports text, image, and audio input, variable reasoning effort, and context windows up to one million tokens. Fine-tuning is available through Tinker.

The company reports more than 80 percent on SWE-bench Verified and prices output at $1.20 per million tokens, versus $4.05 for the larger Inkling. Those are vendor-reported results and include harness choices that matter: the published benchmark table notes a bash-only setup for SWE-bench and an internal coding harness for Terminal-Bench.

The interesting product boundary is not simply “small.” A 12B-active executor can still carry long multimodal context and tool-oriented reasoning while consuming less inference compute. Teams should reproduce results in their own harness, especially for factuality and safety, before treating the open weights as a drop-in frontier replacement.

## Source

- [Thinking Machines](https://thinkingmachines.ai/news/inkling-small)
