Open Multimodal Decision Models Set New Standards for Edge AI Efficiency
The release introduces a new generation of open-weight, multimodal decision models, d1-3B and d1-omni-600M, designed specifically for high-speed, structured decision-making, particularly optimized for deployment on edge devices. These models represent a significant architectural departure from traditional Large Language Models (LLMs) or generative AI systems. Unlike generative models, which operate by predicting and producing sequences of tokens (a process that can be computationally intensive and sequential), these decision models function by executing a single, efficient forward pass. This fundamental difference allows them to provide direct, structured answers to complex queries, making them exceptionally fast and reliable for real-time applications.
The core strength of these models lies in their ability to process multiple data types—multimodality—while maintaining a remarkably small parameter count, allowing them to achieve state-of-the-art performance even on resource-constrained hardware.
Two distinct models are presented: d1-3B and d1-omni-600M. The d1-3B model is built upon the LFM2.5-VL-3B backbone, which is described as a decoder-only Vision-Language Model (VLM). This architecture enables d1-3B to accept and process inputs combining text and images. Meanwhile, the d1-omni-600M model utilizes a different, highly versatile foundation: LFM2.5-Encoder-350M. This backbone is a bidirectional encoder, a type of neural network architecture known for its ability to process context from both directions (left-to-right and right-to-left), which is crucial for deep understanding. Critically, d1-omni-600M is designed for even broader multimodal input, accepting either a combination of text and images, or alternatively, a combination of text and audio. It is noted that d1-omni-600M is currently in an early research release phase, indicating ongoing development and refinement.
To validate their capabilities, the developers benchmarked both d1-3B and d1-omni-600M across seven diverse public datasets. These datasets cover a wide spectrum of real-world AI tasks, including reading comprehension, toxicity detection (identifying harmful content), intent classification (determining the user's goal), medical Question Answering (QA), and cross-lingual understanding (understanding meaning across different languages). The results demonstrate superior performance: d1-3B achieved a mean score of 82.9, establishing the highest score in the tested comparison and surpassing even the performance of the Decider 4B model. d1-omni-600M scored 78.4, which is highly impressive as it surpasses the Decider 2B model's score (77.1) while containing only a quarter of the parameters. This efficiency gain—high performance with a minimal footprint—is a major selling point for edge computing.
Beyond general benchmarks, the models were validated for their specific multimodal retention. It was confirmed that d1-3B retains the robust vision capabilities inherent in its LFM2.5-VL-3B backbone when tested on standard vision benchmarks. Similarly, d1-omni-600M successfully handles the integration of all three modalities (text, image, and audio). The report acknowledges that while the Decision Index v0.3 includes a private vision split, comprehensive audio decision benchmarks remain an open research problem, which is a necessary caveat for users.
Deployment and efficiency are central to this release. The models were evaluated in collaboration with NVIDIA across a range of hardware, including the high-end NVIDIA GeForce RTX 4090 GPU and various Jetson embedded systems (AGX Thor, AGX Orin, and Orin Nano). The performance metrics highlight the exceptional speed of d1-3B, especially in edge inference scenarios. On the edge, d1-3B can answer a single question in under 50 milliseconds across all measured Jetson devices. Furthermore, the speed scales efficiently: processing three questions only takes 1.3 times the time required for a single question, with the advanced AGX Thor showing an impressive jump from 16 ms to just 20 ms. When deployed on powerful GPU infrastructure, d1-3B maintains its speed, answering a single question in under 10 milliseconds and processing a 384px image in under 18 milliseconds on both tested GPU platforms. This combination of low latency and high accuracy makes the models ideal for mission-critical, real-time applications.
In summary, the d1 decision models are positioned as the go-to solution when the requirement is for fast, structured, and multimodal decision-making. d1-3B is highlighted for delivering the highest decision quality relative to its size, making it a powerful general-purpose choice. Conversely, d1-omni-600M is specifically recommended for scenarios where the physical footprint and parameter count are the most critical limiting factors. Both models are openly available as open-weight models on Hugging Face, encouraging widespread adoption and research, and the developers have provided necessary code and documentation for immediate use.
Why it matters
- —They fundamentally shift the paradigm from generative token prediction to single-pass decision making, enabling faster, more reliable real-time applications.
- —The models achieve state-of-the-art performance (e.g., d1-3B's 82.9 mean score) while maintaining extremely small parameter counts, making them ideal for resource-constrained edge devices.
- —The multimodal capability (text, image, and audio) combined with low latency makes them highly valuable for complex, real-world sensing and decision systems.
Key facts
- d1-3B and d1-omni-600M are decision models that use a single forward pass, unlike generative models that produce tokens.
- d1-3B achieves the highest mean score (82.9) on seven public multimodal datasets, surpassing competitors like Decider 4B.
- The models demonstrate exceptional edge efficiency, with d1-3B answering a single question in under 50 ms on Jetson AGX Thor.
- d1-omni-600M is a highly efficient multimodal model, scoring 78.4 while using only a quarter of the parameters of a competitor model.
- Both models are open-weight and available on Hugging Face for immediate use.
The full text is in the original source. Here we provide a brief summary and key facts.