Thinking Machines LabOpen Source

Inkling

Compare this model

Inkling is Thinking Machines Lab's first train-from-scratch model, released July 15, 2026 under the Apache 2.0 license. A 975B-parameter sparse Mixture-of-Experts (41B active per token) with a 1,000,000-token context window, it is the leading U.S. open-weights model on the Artificial Analysis Intelligence Index (score 41). It accepts text, image, and audio input and generates text, trained on ~45 trillion tokens across text, image, audio, and video. Positioned as a customizable fine-tuning base rather than a finished chat product, it scores 77.6% on SWE-Bench Verified and 97.1% on AIME 2026. Weights are free to self-host; the Tinker API runs $1.87/$4.68 (64K) to $3.74/$9.36 (256K) per million tokens.

Parameters

Undisclosed

Context Window

1M

License

Apache 2.0

Release Date

2026-07-15

Japanese Language Capability

🌐Multilingual

General multilingual model. Basic Japanese processing is possible, but inferior to specialized models.

API Pricing

Input Price (per 1M tokens)

$3.74

Output Price (per 1M tokens)

$9.36

Billing Mode: standard

Strengths

  • Strong logical reasoning and problem-solving capabilities.
  • Designed for structured thinking and analysis tasks.
  • Effective at breaking down complex problems into steps.

Weaknesses

  • May lack extensive general knowledge outside reasoning domains.
  • Performance on creative or open-ended tasks might be limited.
  • Dependent on clear input for optimal reasoning chains.

Use Cases

  • Complex problem solving and logical deduction tasks.
  • Mathematical proofs and technical analysis.
  • Workflow planning and strategic decision support.

Deep Analysis

Total/Active Parameters

975B/41B

Mixture-of-Experts architecture

Context Window

1M tokens

Available via open weights

SWE-Bench Verified

77.6%

vs. Nemotron 3 Ultra: 70.7%

GPQA Diamond

87.2%

Strong science reasoning

VoiceBench

91.4%

Top-tier open-weights audio performance

Input Price (64K ctx)

$1.87/1M

Via Tinker API

FORTRESS Adversarial

78.0%

Best open-weights safety

Strengths

  • Native multimodal reasoning over text, image, and audio with a unified architecture
  • Controllable thinking effort (0.2-0.99) for precise cost-performance optimization
  • Strong agentic and coding capabilities with excellent token efficiency
  • Apache 2.0 license enabling unrestricted enterprise customization and deployment
  • Robust safety profile with best-in-class refusal of harmful queries among open models

Weaknesses

  • Factuality scores (SimpleQA: 43.9%) significantly trail leading closed models like GPT-5.6 Sol (71.6%)
  • Massive hardware requirements (2TB VRAM for BF16) limit accessible local deployment
  • Peak reasoning (HLE text-only: 29.7%) is outperformed by specialized models like GLM 5.2 (40.1%)
  • Creative writing and nuanced instruction following lag behind frontier closed models

Competitor Comparison

ModelArenaSWEGPQAPrice
GLM 5.2N/A80.0%89.5%Open Weights
DeepSeek V4 ProN/A80.6%88.8%Open Weights
Claude Fable 5 (max)N/A95.0%92.6%Closed Weights
GPT-5.6 Sol (xhigh)N/A82.2%94.1%Closed Weights

Inkling is Thinking Machines Lab's inaugural open-weights foundation model, designed as a broad, balanced generalist optimized for customization rather than benchmark dominance. With 975B total parameters (41B active) via a sparse Mixture-of-Experts architecture, it natively processes text, images, and audio through a unified encoder-free design. The model's key innovation is its controllable thinking effort, allowing developers to dial the reasoning budget from 0.2 to 0.99 to precisely balance performance against token cost and latency.

Positioned as a practical multimodal foundation for enterprise fine-tuning, Inkling achieves competitive but not state-of-the-art performance across reasoning, coding, and safety benchmarks while offering unique advantages in native multimodality, token efficiency, and true open-source licensing (Apache 2.0). It excels in agentic workflows and represents the leading open-weights model from a U.S. lab, though it trails specialized Chinese open models on pure reasoning and top closed models on peak capabilities.

The release signals Thinking Machines' focus on building customizable, efficient, and trustworthy AI systems rather than pursuing frontier scale alone. The model is available for immediate fine-tuning via the company's Tinker platform, with full weights on HuggingFace and partnerships with major inference providers for deployment.

Analysis generated: 2026-07-17