Companies

Moondream

moondream.ai

Moondream builds efficient vision-language models for real-time visual AI across cloud, edge, and on-device deployments.

HQSan Francisco, California, United States
Employees11-50
Funding$4.5M
6 active roles
Profile 6mo agoJobs checked 21h ago
AI / MLFoundation Model ProviderB2B SaaSPre-Seed$1M-$10M

About

Moondream builds open-source, deployment-friendly vision-language models and supporting fine-tuning and inference software for developers and organizations deploying visual AI in cloud, edge, and on-device settings. Its differentiation is efficient, low-latency visual reasoning that can run in real time across laptops, GPUs, Jetson devices, and production systems.

Market

Moondream competes in open-source vision-language models and production visual AI, especially compact, low-latency models for edge, embedded, private-cloud, and real-time applications. It differentiates from broader multimodal chatbot models through native grounded outputs, production-oriented skills, small and quantizable models, open weights, and fine-tuning for specific customer workflows. Its cloud-to-edge model, common API, and Photon runtime let teams start with hosted inference and move to local, on-premises, or air-gapped deployment without changing the model interface.

Target Customers

Moondream targets developers and ML/AI or product-engineering teams—from startups and individual builders to enterprises—that need production visual AI for robotics, autonomous agents, security and surveillance, document intelligence, smart automation, or other image and camera workflows. Its strongest fit is for buyers who need low-latency, privacy-sensitive, resource-constrained, or air-gapped deployment and the ability to fine-tune models on proprietary data, while retaining a managed cloud option.

At a Glance

Problem

Moondream addresses the production gap in vision-language AI: larger frontier models can be slow, difficult to run locally, and unreliable on the “last mile” of a company’s specific visual data. For production systems, sending images to the cloud can also introduce latency, recurring inference costs, privacy exposure, and dependence on network connectivity. The core pain is therefore not merely understanding an image in a demo, but making accurate visual decisions quickly and economically in the real world.

Its clearest use case is real-time visual perception at the edge—for robotics, industrial inspection, cameras, kiosks, vehicles, and embedded products. These applications need local, low-latency detection, OCR, grounding, or classification, often in settings where cloud round trips are too slow, too costly, or unacceptable for security and privacy reasons.

Product / Service

Moondream provides a production-oriented visual AI stack built around small, open-weight vision-language models. Its model lineup ranges from a 9B mixture-of-experts Moondream 3 Preview with 2B active parameters, to a 2B production-stable model and a 0.5B model designed for highly constrained devices. The models support tasks such as image querying, detection, pointing, captioning, OCR, segmentation, structured output, and visual reasoning, while permissive licensing allows personal, research, and most commercial use.

The company offers several delivery modes that share the same general model and API experience: hosted inference through Moondream Cloud, local and air-gapped deployment through its Photon inference engine, and Lens for supervised or reinforcement-learning fine-tuning on a customer’s data. This lets developers start in the cloud and move to local hardware without rebuilding the application. Cloud and Lens are monetized on a pay-per-token basis, while the models and Photon are free; Moondream reports 20-millisecond inference on an H100 and Cloud pricing designed to be cheaper than comparable general-purpose multimodal APIs for grounding-heavy workloads.

Market

Moondream competes in the efficient vision-language-model and edge visual-AI market, spanning open-source model weights, hosted multimodal APIs, fine-tuning platforms, and inference infrastructure. Its direct hosted alternatives include vision-capable general-purpose APIs such as Google’s Gemini 2.5 Flash and OpenAI’s GPT-5 Mini, which Moondream compares against on runtime, token consumption, and cost. In the open-model segment, it sits alongside smaller vision-language models such as Qwen2.5-VL and SmolVLM, while differentiating through an integrated model, fine-tuning, inference, and deployment stack rather than selling weights alone.

Moondream is commercializing rather than remaining pre-revenue: it launched Moondream Cloud in 2025, charges for Cloud and Lens usage, and offers paid team and enterprise plans with on-premises deployment, compliance options, and support. Public traction includes more than 5 million monthly downloads, a reported $4.5 million 2024 pre-seed round from Felicis, M12, and Ascend, and a July 2026 partnership that made Moondream 3.1 available through Cloudflare Workers AI. Public customer counts and revenue totals are not disclosed, so the strongest visible indicators are developer adoption, open-model distribution, paid product availability, and ecosystem partnerships rather than reported financial scale.

Founders & Leadership

Jay AllenFounder
CEO
Vik KorrapatiFounder
CTO
Conway AndersonProduct & Design Leader

Funding History

2024-10
Pre-Seed$4.5M

Felicis

Recent News

2026-07-07partnership
Moondream 3.1: Beyond Benchmarks

Moondream launched version 3.1 with best-in-class benchmark scores, a fine-tuning approach designed to transfer to customer tasks, and a partnership with Cloudflare.

2026-05-01product
Photon 1.2.0: Faster Inference, Now on Mac, Windows, Blackwell, and Jetson Thor

Photon 1.2.0 added native inference on Apple Silicon and Windows, support for NVIDIA Blackwell and Jetson Thor, and speed improvements across existing GPUs.

2026-04-20product
Lens: Moondream's Finetune Service

Moondream announced Lens, a fine-tuning product intended to solve the last-mile problem and make vision-language models production-ready.

2026-04-16
moondream - Vision-Language Models

MathWorks documentation described using Moondream to generate descriptive image captions, highlighting the model's lightweight design and speed.

2026-03-25product
Photon: Real-Time VLM Is Here

Moondream announced Photon for real-time production vision AI, targeting deployments ranging from edge devices to H100-class servers. The update also highlighted 4-bit quantization for faster, smaller inference.

2026-03-10product
Moondream Segmenting Update: Better Masks, Better Benchmarks, 40% Faster

Moondream upgraded its Cloud segmentation capability with better mask quality, stronger benchmark results, and inference that is 40% faster than before.

2026-02-19partnership
PTZOptics Launches Visual Reasoning Initiative, Partners with Moondream to Automate Video Decision-Making

PTZOptics announced a partnership with Moondream for a visual-reasoning initiative aimed at automating video decision-making. Moondream supplies an open, lightweight vision model for fast and accurate visual reasoning.

2025-12-19product
We added Moondream 3 Preview support to Moondream Station

Moondream Station added support for Moondream 3 Preview, allowing Mac users to run the model with MLX-native, quantized performance.

2025-10-17product
Announcing Moondream Cloud

Moondream launched Moondream Cloud, a hosted version of its vision model designed to simplify development of vision applications while emphasizing speed and cost efficiency.

2025-09-18product
Moondream 3 Preview: Frontier-level reasoning at a blazing speed

Moondream released a preview of its 9B-parameter mixture-of-experts model with 2B active parameters, frontier-level visual reasoning, and a 32K-token context length. The model was made available through the Moondream playground and Hugging Face.

Active Roles

6
San Francisco, CA/Sales/149d ago
San Francisco, CA (in-person)/Marketing/182d ago
San Francisco, CA · In-person · Full-time/Engineering/182d ago
San Francisco, CA · In-person · Full-time/Design/182d ago
San Francisco, CA · In-person · Full-time/Engineering/182d ago
San Francisco, CA · In-person · Full-time/Operations/182d ago

Business Model

Moondream offers its models and Photon inference engine for free, while monetizing Cloud and Lens through pay-per-token pricing. It also offers paid team and enterprise-oriented plans and services.

Products

Moondream open model family: Moondream 3 Preview/3.1, Moondream 2, and Moondream 2 0.5BLens hosted fine-tuning APIPhoton high-performance inference runtimeMoondream Cloud hosted inference APIMoondream Station for free local model execution

Customers

FashnAIWarmHubQualcomm

Tech Stack

Vision-language models (VLMs) for grounded visual reasoningSparse mixture-of-experts architecture: Moondream 3 uses 9B total parameters with 2B active, 64 experts, and 8 experts activated per tokenNative grounded vision skills for object detection, pointing, captioning, visual question answering, OCR, and segmentation, with structured boxes, coordinates, and masksSupervised fine-tuning and reinforcement learning through the Lens APIPhoton model-specific inference runtime with streaming, automatic batching, prefix caching, and paged KV cacheCUDA/NVIDIA GPU and Jetson deployment across cloud, desktop, server, edge, and private infrastructureOpen-source Apache-2.0 model distribution with quantized fp16, int8, and int4 variants

Competitors

Hugging Face SmolVLM/SmolVLM2
Alibaba Qwen2.5-VL
Google Gemma 3

Key Investors

Felicis, Ascend, M12 GitHub Fund (Microsoft)