Companies

Lamb Labs

lamb-labs.com

Lamb Labs designs and sells ultra-low-power chips for fast AI inference on edge devices.

HQLondon, Not applicable, United Kingdom
Jobs checked 17h ago
AI / MLAI InfrastructureInfrastructure

About

Lamb Labs builds custom silicon for AI inference and a post-training technique that converts existing transformer models into diffusion architectures, enabling parallel decoding. It targets neoclouds, enterprises, trading firms, and edge-device markets; its differentiator is pairing model redesign with purpose-built hardware, with stated goals of 20,000+ tokens per second and 63× higher intelligence per watt than an RTX 6000 Ada GPU.

Market

Lamb Labs competes in AI-inference semiconductors and accelerators, spanning datacenter inference for neoclouds and enterprises as well as ultra-low-power edge inference. Its differentiation is to convert existing transformer models into diffusion architectures that decode tokens in parallel, then hard-code that architecture into custom silicon; the company targets 20,000+ tokens per second and 63× higher intelligence per watt than an RTX 6000 Ada GPU.

Target Customers

Lamb Labs targets technical teams at neoclouds, AI-intensive enterprises, and trading firms seeking lower inference cost or latency. It also targets edge-device companies building robots, AI cameras, wearables, and smart-home products whose models are constrained by power budgets; company size and specific buyer titles are not stated.

At a Glance

Problem

Lamb Labs addresses the growing gap between AI model capability and the compute, power, latency, and connectivity constraints of the devices and infrastructure that must run those models. The company argues that AI’s bottleneck is increasingly compute and power: data-center capacity is expanding more slowly than demand, models are getting larger, and existing hardware and software are not designed together. The economics are therefore lower inference costs and energy use for neoclouds and enterprises, lower latency for trading firms, and practical local inference for devices whose power budgets cannot support larger models.

The clearest use case is private, real-time AI inference at the edge—in robots, AI cameras, wearables, smart-home devices, vehicles, medical and security hardware, and industrial equipment. These applications benefit from keeping data local while avoiding cloud latency, connectivity dependence, and the cost and power demands of sending every inference to a data center.

Product / Service

Lamb Labs combines model transformation with purpose-built inference hardware. Its post-training technique converts an existing autoregressive transformer into a diffusion architecture that decodes multiple tokens in parallel, without a new pretraining run or reported loss in quality. The company says the method works across LLMs and vision-language models and can deliver roughly 2× faster inference on GPUs customers already own, providing a software-level entry point before custom silicon is available.

The longer-term product is custom inference silicon designed around that parallel-decoding architecture. Lamb Labs says it designs and sells chips for on-edge AI inference, has demonstrated an FPGA implementation running an 8-billion-parameter model under 10 watts, and is developing custom ASICs targeting more than 20,000 tokens per second and 63× higher intelligence per watt than an RTX 6000 Ada GPU. The commercial model therefore appears to be a combination of model optimization and hardware sales, with the ASIC claims clearly presented as targets rather than current production performance.

Market

Lamb Labs operates in AI hardware, specifically inference accelerators and low-power edge-AI chips, with an emphasis on combining model architecture and silicon rather than optimizing only one layer. Its alternatives include established edge platforms such as NVIDIA Jetson and dedicated edge processors such as Hailo’s AI chips; customers may also continue using general-purpose GPUs or cloud inference when power, privacy, latency, and connectivity are less restrictive.

The company is very early-stage. Y Combinator lists it as an active Summer 2026 company founded in 2026, while Lamb Labs’ own roadmap marks its diffusion post-training technique and FPGA accelerator as shipped; its prototype reportedly runs an 8B model under 10 watts. Its LinkedIn page lists a two-person company and says a public demo is forthcoming. The reviewed public materials do not announce customers, revenue, or commercial ASIC availability, so Lamb Labs is best characterized as pre-commercial or at most early-revenue, with technical prototypes and YC backing constituting its visible traction rather than validated commercial scale.

Founders & Leadership

Thomas LanningFounder
Co-Founder and CEO
Niki KotechaFounder
Founder

Funding History

2026-07
Seed (Y Combinator)$500K

Y Combinator

Recent News

2026-07-28
Y Combinator profiles Lamb Labs as a builder of custom AI-inference chips

Y Combinator’s company profile describes Lamb Labs as a 2026 London startup building custom chips for AI inference and a high-speed LLM, targeting up to 20,000+ tokens per second and substantially higher intelligence per watt.

2026-07
Lamb Labs launches from stealth with ultra-low-power AI inference chips

A Lamb Labs LinkedIn update says Niki Kotecha and Thomas Lanning are coming out of stealth and launching the company. The startup is focused on ultra-fast, ultra-low-power inference chips for deployments below 1 watt.

2026-07product
Lamb Labs presents diffusion inference architecture and FPGA prototype

Lamb Labs’ official website describes a post-training method that converts transformer models into parallel-decoding diffusion architectures, claiming 2× faster inference on existing GPUs. It also reports a shipped FPGA accelerator prototype running an 8B-parameter model under 10 watts and outlines a future custom-ASIC roadmap.

Active Roles

0

No active roles right now.

Get notified when they post

Business Model

Lamb Labs’ revenue model is hardware-led: it designs and sells custom, ultra-low-power chips for on-edge AI inference. Its intended B2B customers include neoclouds, enterprises, trading firms, and edge-device makers, although specific pricing and recurring software revenue are not disclosed.

Products

Diffusion post-training/model-conversion technology that enables parallel decoding on existing GPUsFPGA inference accelerator, with an 8B-parameter model prototype running under 10 W on a Kria KV260Merino custom accelerator board with additional memory for large modelsModel-specific custom ASICs for ultra-fast, energy-efficient AI inference

Customers

None publicly disclosed

Tech Stack

Diffusion-model post-training and parallel token decodingTransformer-based LLM and VLM supportFPGA acceleration, currently prototyped on the Kria KV260Model-specific custom ASIC siliconHard-coded compute and memory layoutsReinforcement-learning-based chip-design optimization

Competitors

Groq
Cerebras
d-Matrix
Tenstorrent
Hailo
Untether AI