Multimodal Open D1 Decision Models For The Edge
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Liquid AI released open-weight d1-3B and experimental d1-omni-600M, models designed to return structured decisions in a single forward pass. The company reports strong Decision Index and public-dataset results for d1-3B, plus sub-50-millisecond response times on tested edge devices. The release does not include vision or audio benchmark scores, and no speed results are reported for d1-omni-600M.

Liquid AI released two open-weight decision models, d1-3B and experimental d1-omni-600M, for tasks that classify or otherwise structure answers from text and, depending on the model, image or audio inputs. The company says d1-3B can answer a question in 16 milliseconds on an NVIDIA Jetson AGX Thor and under 50 milliseconds on each of its other tested edge devices, positioning the models for applications that need local, low-latency decisions rather than generated text.

Liquid AI describes the models as decision systems that return an answer in a single forward pass, rather than generating a sequence of tokens like a conventional language model. The release says d1-3B accepts text and images and is based on the company’s LFM2.5-VL-3B vision-language model. The smaller d1-omni-600M is based on LFM2.5-Encoder-350M and adds vision and audio encoders. It accepts text paired with an image or text paired with audio; Liquid AI labels it an early research release under active development.

On the company’s reported Decision Index 0.2.1 comparison, d1-3B scored 48.57. Liquid AI says that result places it ahead of the 4B and 9B models in the comparison and the 35B-A3B Decider, which scored 47.11. The company also tested both models on seven public datasets spanning reading comprehension, toxicity detection, intent classification, medical question answering and cross-lingual understanding. In its table, d1-3B had a mean score of 82.9, while d1-omni-600M scored 78.4; the listed Decider 4B and Decider 2B means were 81.1 and 77.1, respectively.

Liquid AI and NVIDIA measured d1-3B on an NVIDIA RTX 4090 and three Jetson systems, among other tests. For one question, the company reports 16 ms on Jetson AGX Thor, 26 ms on Jetson AGX Orin 64 GB and 50 ms on Jetson Orin Nano. Its results also list 8 ms on RTX 4090 and 9 ms on AMD MI325X. These are company-reported measurements; the release gives test configurations and workload figures, but independent replication is not included. Liquid AI did not provide speed results for the experimental 600M model.

At a glance
announcementWhen: Announced in 2026; both models are avai…
The developmentLiquid AI has made two multimodal d1 decision models available as open weights, including an experimental compact model intended for edge use.

Local Decisions on Small Devices

The announcement targets developers who need a system to make fast, structured choices without sending every request to a remote service or generating a long response. Examples in the release include routing customer-support tickets, scoring urgency and answering questions about images. If the reported latency holds in a particular deployment, a device could classify inputs and return a result quickly even where network access is limited or cloud processing is undesirable.

The models also offer different size and input trade-offs. Liquid AI presents d1-3B as the higher-scoring option in its reported comparisons and d1-omni-600M as a smaller model with text, image and audio capabilities. That may be relevant for embedded systems with constrained memory or power, but the release does not provide power consumption, memory requirements or a full edge-device speed comparison for the 600M model. Model size alone does not establish how well either system will perform in a given product.

For businesses, the distinction between a decision model and a generative assistant is practical: a fixed classification or scoring task may need a concise, predictable output rather than open-ended text. Still, benchmark results do not establish accuracy for a company’s own categories, languages, inputs or safety requirements. Teams would need to test the models on representative data before relying on them for customer-facing or consequential decisions.

Amazon

edge AI decision-making device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the d1 Models Differ

The d1 release builds on Liquid AI’s Liquid Foundation Models, but the two models use different architectures. d1-3B starts from a decoder-only vision-language model and handles text and images. d1-omni-600M starts from a bidirectional encoder and adds separate vision and audio encoders, allowing either image-and-text or audio-and-text input. Both are described as decision models, with outputs produced in one pass rather than through token-by-token generation.

The seven-dataset table gives a mixed picture across tasks rather than a uniform win. d1-3B leads the reported mean, but it does not score highest on every listed dataset: for example, the table gives Decider 4B higher results on MASSIVE intent, BoolQ and XNLI. d1-omni-600M scores above Decider 2B on the reported mean, while individual dataset outcomes vary. The release’s aggregate scores should therefore be read alongside the task-level results.

Liquid AI says it checked whether d1-3B retained vision capabilities from its underlying model, but it does not report vision benchmark scores in this release. It also provides no audio benchmark results for d1-omni-600M. The company says the Decision Index’s next version includes a private vision split and that audio decision benchmarks remain an open problem. Both models are available through Hugging Face, and Liquid AI points users to demos in its System One Arcade space.

““Unlike our generative models, decision models don’t produce tokens but answer in a single forward pass.””

— Liquid AI

Limits of the Reported Results

The performance and latency figures in the announcement are reported by Liquid AI; the supplied material does not describe independent testing. Results may also depend on hardware, software versions, input format and workload. The release provides several device measurements but does not establish how they translate to a specific application or operating environment.

There are no reported vision or audio benchmark scores, and no speed figures for d1-omni-600M. The company says the model is experimental, so its capabilities and performance may change as development continues. The release also does not provide enough information to judge power use, memory footprint, deployment costs, or performance across a broad range of real-world inputs.

While the public-dataset table reports mean scores and task-level results, it does not by itself show whether the models are suitable for high-impact decisions, how they handle ambiguous cases, or how performance varies across populations and languages. Those questions require additional evaluations and testing by prospective users.

Availability and Further Evaluation

Both models are available now as open weights on Hugging Face, according to Liquid AI, which also directs developers to demonstrations in its System One Arcade Hugging Face Space. The company’s example usage requires Transformers version 5.14 or later and uses model-provided code. Developers considering deployment will need to review the model cards, test latency on their target hardware and evaluate outputs on their own tasks.

Further evidence to watch for includes updated results for the experimental d1-omni-600M, measured inference speed for that model, and published vision or audio evaluations. Liquid AI has not specified a date for those additions. Until such information is available, the release establishes availability and company-reported text and general decision benchmarks, but leaves important multimodal and deployment questions open.

Key Questions

What did Liquid AI release?

Liquid AI released d1-3B and d1-omni-600M as open-weight decision models. The first handles text and images; the second is an experimental model that accepts text paired with an image or audio.

What does a decision model do?

Liquid AI says its decision models return structured answers in a single forward pass, rather than generating a sequence of tokens. The company’s examples include categorizing support requests and scoring urgency.

How fast is d1-3B on edge hardware?

Liquid AI reports one-question response times of 16 ms on Jetson AGX Thor, 26 ms on Jetson AGX Orin 64 GB and 50 ms on Jetson Orin Nano. These are company-reported measurements, not independently verified results in the release.

Are vision and audio capabilities benchmarked?

No numerical vision or audio benchmark results are included. Liquid AI says its reported public-dataset evaluation covers seven other tasks and that it is not reporting vision or audio scores in this release.

Where can developers get the models?

Liquid AI says both open-weight models are available on Hugging Face. The company also links to demos in its System One Arcade space; its usage example specifies Transformers 5.14 or later.

Source: rss

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI Tower: Twelve Rooms Of Safe AI Workspaces In Action

Exploring the AI Tower’s twelve rooms, a new framework for safe and effective AI workspaces, now in practical use for organizations and developers.

How A Tiny Anthropic Team Became A Surprisingly Productive Hit Factory

A small team at Anthropic has become unexpectedly prolific in AI development, raising questions about team size and innovation efficiency.

How To Access Grok Bot With Expanded X.ai Subscription Options

xAI has announced broader access to Grok Bot across more subscription plans, lowering barriers for users to utilize its AI assistant features.

Retrospectively Reverse-Engineering Apple’s Neural Engine

Experts are now analyzing Apple’s Neural Engine through reverse-engineering efforts, revealing new insights into its design and capabilities.