Co-Founder & Lead Programmer of AcceleratedLogic AI
Thinking Machines Lab released Inkling on July 15, 2026 as an open-weights model. The model is large even though it uses a Mixture-of-Experts architecture, so the active parameter count should not be mistaken for its full memory requirement. This guide summarizes the published model card and explains the practical hardware and license details to check before planning a deployment.
Model size, inputs, and license
Thinking Machines' Inkling model card lists 975 billion total parameters and 41 billion active parameters, a context window up to one million tokens, and text, image, and audio inputs with text output. The card describes a 66-layer multimodal mixture-of-experts transformer and lists the license as Apache 2.0. Its release announcement says the model was pretrained on text, images, audio, and video.
Active parameters describe which experts are used for a token; they do not tell you the disk size or total memory needed to keep the complete model available. That distinction matters when comparing a sparse model with a smaller dense checkpoint.
Hardware requirements are substantial
The current model card says the BF16 checkpoint requires at least 2 TB of aggregated GPU VRAM. It lists an NVFP4 alternative at at least 600 GB of aggregated VRAM, with specified multi-GPU configurations. Those are cluster-scale requirements, not typical laptop or single-GPU desktop requirements. Hardware and quantization guidance can change, so use the provider's current card and serving documentation when planning infrastructure.
Open weights allow a developer to download and run the checkpoint under its license, but they do not make deployment cheap or remove the need to review model terms, hardware requirements, data handling, and any third-party serving provider. If you call Inkling through an external API, that request is still processed by that provider.
Inkling-Small is a separate option
Thinking Machines also announced Inkling-Small as a lighter-weight sibling. Its current model page lists different parameter counts and context limits for provider-hosted Tinker access than for the full checkpoint. Do not assume the two sizes have identical costs, limits, or local hardware needs; inspect the card for the exact model and route you plan to use.
How to evaluate it before adopting it
For a local deployment, estimate memory for the full checkpoint and its runtime overhead, not only the active parameter count. Include the context length and concurrent requests in the estimate. Verify that your inference framework supports the selected quantization and that the license fits the way you plan to distribute or host the system.
For quality, define a small set of representative coding, retrieval, and multimodal tasks. Run the same prompts and evaluation harness across Inkling, Inkling-Small, and your current model. Track task success, latency, memory use, and failure modes separately. Provider descriptions such as “well-rounded” or “open weights” are useful context, but they are not a substitute for measurements on your workload.