🔍 Read the full analysis: Can Open D1 Decision Models Process Multiple Modalities At The Edge? on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Liquid AI has released two open-weight decision models designed to return structured answers in a single forward pass, including models that accept image inputs and, experimentally, audio. The company reports strong results on selected text benchmarks and sub-50-millisecond d1-3B responses on tested devices, but independent evaluations and published vision or audio benchmark results are absent.
Liquid AI has released d1-3B and d1-omni-600M, open-weight models built to classify, score and answer decision tasks in a single forward pass. The company says the models support combinations of text with images, and that the smaller experimental model also supports text with audio, putting them forward for applications where latency and edge-device constraints matter; independent multimodal results have not been published.
Liquid AI says d1-3B is based on its LFM2.5-VL-3B vision-language model and accepts text and images. The company reports a score of 48.57 on Decision Index 0.2.1. It also reports that the model answered one question in 16 milliseconds on an NVIDIA Jetson AGX Thor, 26 milliseconds on a Jetson AGX Orin 64 GB and 50 milliseconds on a Jetson Orin Nano. Those are company measurements, not independently replicated results in the material provided.
The smaller d1-omni-600M uses the LFM2.5-Encoder-350M bidirectional encoder with added vision and audio encoders. Liquid AI describes it as an early research release still under development, supporting text paired with an image or audio. The company reports no speed measurements for this model, and it has not published vision or audio benchmark scores for either release.
On seven public datasets focused on text tasks—including reading comprehension, toxicity detection, intent classification, medical question answering and cross-lingual understanding—Liquid AI reports mean scores of 82.9 for d1-3B and 78.4 for d1-omni-600M. Its comparison table lists Decider 4B at 81.1 and Decider 2B at 77.1, respectively. Results vary by dataset: d1-3B scores below Decider 4B on BoolQ, MASSIVE intent and XNLI. The company also reports that three questions took 1.3 times as long as one on tested devices; on Jetson AGX Thor, it gives 16 milliseconds for one and 20 milliseconds for three.
Decision Models at the Edge
The release targets a different use from a general-purpose text generator: a model could return a structured decision such as which team should receive a customer request, how urgent it is, or an answer to a question about an image. If its speed and accuracy hold in a particular deployment, that approach may suit applications that need a quick classification without generating a longer response.
Edge processing can matter where sending data to a remote service adds delay, connectivity dependence or data-handling concerns. Liquid AI’s reported Jetson timings make the models relevant to developers evaluating local inference on compact hardware. But the measurements describe selected tests, not an assurance of performance in production. Input length, software configuration, task design and device can all affect results, and the release does not establish error rates or the need for human review.
The reported smaller-model benchmark result may also interest teams trying to fit decision workloads into tighter compute budgets. The comparison is limited to Liquid AI’s selected datasets and evaluation; it does not show that the model will outperform alternatives across other tasks, or that its multimodal capabilities match its text results.
As an affiliate, we earn on qualifying purchases.
How the Models Differ
Liquid AI presents both releases as decision models built on its Liquid Foundation Models. Rather than producing a sequence of generated tokens as a conventional chat response, they are intended to return structured outputs in one forward pass. That design is aimed at bounded tasks such as routing, scoring and classification, although the release does not detail how every output format is configured.
The larger model’s underlying vision-language system supports text and images. The smaller omni model combines a bidirectional text encoder with vision and audio encoders, but Liquid AI labels it experimental. The distinction matters: its broader input options do not yet come with published audio or vision decision scores in this release.
The company’s text evaluation covers SQuAD 2.0, Civil Comments, MASSIVE intent, PubMedQA, BoolQ, XNLI and PAWS-X. Liquid AI says Decision Index version 0.3 has only a private vision split and that audio decision benchmarks remain an open problem. These limits mean the public scores describe a specific selection of text-focused tests, not a complete assessment of multimodal decision quality.
““Best decision model under 10B on the Decision Index 0.2.1””
— Liquid AI
As an affiliate, we earn on qualifying purchases.
Limits of the Published Results
The release material does not include independent evaluations, confidence intervals or enough methodological detail to determine how closely the tests match a specific deployment. The reported averages cover seven public datasets, and individual results are not uniformly higher than the comparison model. They cannot establish accuracy or reliability across every decision task.
For multimodal use, the main evidence gap is direct evaluation: Liquid AI says d1-3B retains vision capabilities and d1-omni-600M handles its supported modalities, but publishes no vision or audio benchmark scores. It also reports no speed tests for d1-omni-600M. The source material does not explain how either model handles ambiguous inputs, how often decisions might require human review, or how performance changes across varied production workloads.
The reported response times also come from specified tests on selected hardware. They do not establish a general latency guarantee or show how the models perform with different inputs, software versions or application requirements. Those questions require testing in the intended environment.
As an affiliate, we earn on qualifying purchases.
Developer Testing and Evaluation
Both models are available as open weights on Hugging Face, and Liquid AI points users to demos in its System One Arcade Hugging Face Space. Its release instructions specify Transformers version 5.14 or later and say to load the models with their supplied code enabled.
The next useful evidence will come from developers testing the models against their own tasks and hardware, alongside further published evaluation of image and audio decisions. Until such results are available, the company’s scores and timings should be treated as reported measurements with a defined scope, not proof of performance across edge deployments. Liquid AI has not specified in the source material when expanded benchmarks or a further update to the experimental omni model will arrive.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did Liquid AI release?
Liquid AI released d1-3B and d1-omni-600M, open-weight models intended to return structured answers for decision tasks in a single forward pass.
Can the models process more than text?
According to Liquid AI, d1-3B accepts text and images. The experimental d1-omni-600M supports text paired with an image or audio. The release provides no vision or audio benchmark scores.
How fast is d1-3B on edge devices?
Liquid AI reports one-question response times of 16 milliseconds on Jetson AGX Thor, 26 milliseconds on Jetson AGX Orin 64 GB and 50 milliseconds on Jetson Orin Nano. These are company-reported test results, not independently verified deployment guarantees.
Are the benchmark results independent?
Not in the source material. Liquid AI reports mean scores from seven public text datasets, but independent results and confidence intervals are not provided, and scores vary across individual datasets.
Where can developers access the models?
Liquid AI says both are available as open weights on Hugging Face and points to demos in its System One Arcade Hugging Face Space. Its instructions specify Transformers 5.14 or later and use of the supplied code when loading the models.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
