📊 Full opportunity report: Why Inkling Is The Most Exciting AI Development Right Now on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Thinking Machines has launched Inkling, a massive 975-billion-parameter multimodal AI model, on Hugging Face. While its scale and open availability are notable, performance benchmarks and licensing details are still unclear, limiting immediate practical use.
Thinking Machines has released Inkling on Hugging Face, a 975-billion-parameter multimodal model designed for processing text, images, and audio within a claimed one-million-token context window. The release is significant because it pairs a record scale with open access, although the model’s demanding hardware requirements and limited independent evaluation performance benchmarks and licensing details temper immediate adoption.
The model, described as a decoder-only Mixture-of-Experts architecture, was trained on 45 trillion tokens across multiple data types, including text, images, and audio, according to Hugging Face. It features 256 experts, with six routed experts active during processing, and employs a combination of global and sliding-window attention mechanisms. Its architecture supports hierarchical image patching and audio converted into mel-spectrograms, aiming for versatile multimodal reasoning.
However, the release lacks detailed benchmark results, safety evaluations, and licensing specifics. The hardware requirements are substantial: the BF16 checkpoint demands approximately 2 TB of VRAM, and the NVFP4 version requires about 600 GB, making full deployment impractical for most users without specialized infrastructure. Developers can access the model via hosted inference services or run it on supported frameworks like Transformers, SGLang, vLLM, and llama.cpp, but real-world performance remains unverified.
Implications of Inkling’s Open Release for AI Development
Inkling’s release marks a major milestone in scaling multimodal AI models, offering a new tool for research and enterprise applications that require integrated processing of language, images, and audio. Its open access could accelerate innovation in scientific research, media analysis, and complex data workflows. However, the high hardware demands and lack of independent performance benchmarks mean its practical impact will depend on future evaluations and community testing.

PNY NVIDIA GeForce RTX™ 5060 Ti OC Dual Fan, Graphics Card (16GB GDDR7, 128-bit, Boost Speed: 2692 MHz, SFF-Ready, PCIe® 5.0, HDMI®/DP 2.1, 2-Slot, NVIDIA Blackwell Architecture, DLSS 4)
- AI-Enhanced Rendering: DLSS boosts FPS and image quality
- Advanced GPU Cores: Fifth-Gen Tensor, New Streaming, Ray Tracing Cores
- Optimized Gaming Response: Reflex tech for faster reactions and precision
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on Large-Scale Multimodal AI Models
Previous multimodal models, such as OpenAI’s GPT-4 and Meta’s ImageBind, have demonstrated the potential for integrated reasoning across different data types but often remain proprietary or limited in scale. Inkling’s 975-billion-parameter size surpasses many existing models, positioning it as one of the largest open multimodal architectures. Its development reflects ongoing industry trends toward larger, more versatile models capable of handling complex, multi-data inputs, with open availability potentially democratizing access to such advanced AI tools.
“This model is huge.”
— Hugging Face
Unverified Aspects of Inkling’s Performance and Licensing
It remains unclear how Inkling performs across real-world multimodal tasks, especially video processing, as native video evaluation has not been conducted. The licensing terms and whether the model’s training data and code are publicly available are also unspecified. Additionally, the impact of the model’s speculative prediction layers on inference speed and accuracy is yet to be demonstrated through independent testing.
Next Steps for Evaluating and Using Inkling
Developers and research organizations will likely begin testing Inkling via supported inference frameworks, focusing on latency, accuracy, and hardware feasibility. Independent benchmarks and safety assessments are expected to follow, alongside potential fine-tuning for specific domains. The community’s evaluation will determine whether Inkling’s scale translates into practical advantages or remains primarily a technological milestone.
Key Questions
What is Inkling?
Inkling is a multimodal AI model from Thinking Machines, featuring 975 billion parameters, capable of processing text, images, and audio within a large context window. It is designed for advanced reasoning across multiple data types.
Can I run Inkling on my personal computer?
No. The hardware requirements are extremely high—around 2 TB of VRAM for BF16 and 600 GB for NVFP4 checkpoints—making it impractical for typical consumer systems. Most users will access it via hosted inference services.
Does Inkling support video processing?
While the architecture includes inputs with a temporal dimension that could support video, native video performance has not been evaluated or confirmed at this stage.
What are the licensing terms for Inkling?
The release describes Inkling as an open model but does not specify licensing details, usage restrictions, or whether training data and code are publicly available. Further disclosures are expected.
What are the next steps for Inkling’s development?
Expectations include independent performance testing, safety evaluations, and potential domain-specific fine-tuning. Community feedback and benchmarking will shape its future adoption and applications.
Source: ThorstenMeyerAI.com