Inkling is Thinking Machines Lab's first train-from-scratch model, released July 15, 2026 under the Apache 2.0 license. A 975B-parameter sparse Mixture-of-Experts (41B active per token) with a 1,000,000-token context window, it is the leading U.S. open-weights model on the Artificial Analysis Intelligence Index (score 41). It accepts text, image, and audio input and generates text, trained on ~45 trillion tokens across text, image, audio, and video. Positioned as a customizable fine-tuning base rather than a finished chat product, it scores 77.6% on SWE-Bench Verified and 97.1% on AIME 2026. Weights are free to self-host; the Tinker API runs $1.87/$4.68 (64K) to $3.74/$9.36 (256K) per million tokens.
Parameters
Undisclosed
Context Window
1M
License
Apache 2.0
Release Date
2026-07-15
Japanese Language Capability
General multilingual model. Basic Japanese processing is possible, but inferior to specialized models.
API Pricing
Input Price (per 1M tokens)
$3.74
Output Price (per 1M tokens)
$9.36
Billing Mode: standard
Strengths
- •Strong logical reasoning and problem-solving capabilities.
- •Designed for structured thinking and analysis tasks.
- •Effective at breaking down complex problems into steps.
Weaknesses
- •May lack extensive general knowledge outside reasoning domains.
- •Performance on creative or open-ended tasks might be limited.
- •Dependent on clear input for optimal reasoning chains.
Use Cases
- •Complex problem solving and logical deduction tasks.
- •Mathematical proofs and technical analysis.
- •Workflow planning and strategic decision support.
Deep Analysis
Total/Active Parameters
975B/41B
Mixture-of-Experts architecture
Context Window
1M tokens
Available via open weights
SWE-Bench Verified
77.6%
vs. Nemotron 3 Ultra: 70.7%
GPQA Diamond
87.2%
Strong science reasoning
VoiceBench
91.4%
Top-tier open-weights audio performance
Input Price (64K ctx)
$1.87/1M
Via Tinker API
FORTRESS Adversarial
78.0%
Best open-weights safety
Strengths
- ・Native multimodal reasoning over text, image, and audio with a unified architecture
- ・Controllable thinking effort (0.2-0.99) for precise cost-performance optimization
- ・Strong agentic and coding capabilities with excellent token efficiency
- ・Apache 2.0 license enabling unrestricted enterprise customization and deployment
- ・Robust safety profile with best-in-class refusal of harmful queries among open models
Weaknesses
- ・Factuality scores (SimpleQA: 43.9%) significantly trail leading closed models like GPT-5.6 Sol (71.6%)
- ・Massive hardware requirements (2TB VRAM for BF16) limit accessible local deployment
- ・Peak reasoning (HLE text-only: 29.7%) is outperformed by specialized models like GLM 5.2 (40.1%)
- ・Creative writing and nuanced instruction following lag behind frontier closed models
Competitor Comparison
| Model | Arena | SWE | GPQA | Price |
|---|---|---|---|---|
| GLM 5.2 | N/A | 80.0% | 89.5% | Open Weights |
| DeepSeek V4 Pro | N/A | 80.6% | 88.8% | Open Weights |
| Claude Fable 5 (max) | N/A | 95.0% | 92.6% | Closed Weights |
| GPT-5.6 Sol (xhigh) | N/A | 82.2% | 94.1% | Closed Weights |
Inkling is Thinking Machines Lab's inaugural open-weights foundation model, designed as a broad, balanced generalist optimized for customization rather than benchmark dominance. With 975B total parameters (41B active) via a sparse Mixture-of-Experts architecture, it natively processes text, images, and audio through a unified encoder-free design. The model's key innovation is its controllable thinking effort, allowing developers to dial the reasoning budget from 0.2 to 0.99 to precisely balance performance against token cost and latency.
Positioned as a practical multimodal foundation for enterprise fine-tuning, Inkling achieves competitive but not state-of-the-art performance across reasoning, coding, and safety benchmarks while offering unique advantages in native multimodality, token efficiency, and true open-source licensing (Apache 2.0). It excels in agentic workflows and represents the leading open-weights model from a U.S. lab, though it trails specialized Chinese open models on pure reasoning and top closed models on peak capabilities.
The release signals Thinking Machines' focus on building customizable, efficient, and trustworthy AI systems rather than pursuing frontier scale alone. The model is available for immediate fine-tuning via the company's Tinker platform, with full weights on HuggingFace and partnerships with major inference providers for deployment.
Sources
- Introducing Inkling - Thinking Machines Official Announcement
- Inkling Model Card - Thinking Machines
- Inkling on Hugging Face - Model Repository
- Welcome Inkling by Thinking Machines - Hugging Face Blog
- Thinking Machines has released Inkling - Artificial Analysis
- Thinking Machines open sources Inkling - VentureBeat
Analysis generated: 2026-07-17