Companies

Infinity

infinity.inc

Infinity generates production-ready AI inference software for new chips, enabling hardware partners to deploy models quickly across modalities.

HQSan Francisco, California, United States
Employees11-50
Funding$15M
Valuation$100M
1 active role
Profile 1mo agoJobs checked 17h ago
AI / MLAI InfrastructureB2B SaaSSeries A$10M-$50M

About

Infinity is an AI infrastructure company building inference software that turns hardware specifications into production-ready libraries and full inference stacks for new AI chips. It sells to chip designers and hardware partners, differentiating through rapid, universal enablement across accelerators and modalities, with its Ignition research agent generating inference stacks for any chip.

Market

Infinity competes in AI inference infrastructure and accelerator software, positioning itself as a hardware-independent software layer that delivers Day 1 model enablement and a complete production inference stack for new silicon. Its differentiation is the autonomous generation and optimization of low-level kernels and full-model runtimes in days rather than the months or years typically associated with building chip-specific software, whereas major alternatives such as TensorRT and ROCm are primarily vendor ecosystems and Modular provides a portable multi-hardware platform.

Target Customers

Infinity primarily targets AI accelerator and semiconductor companies—especially new-silicon designers and hardware platform vendors—that need production-quality model support quickly. The likely buyers are CTOs, VPs of engineering, chip-architecture teams, and ML-systems leaders responsible for software enablement and inference performance.

At a Glance

Problem

AI-chip companies can design silicon with impressive theoretical performance yet struggle to make it useful for production inference because the required kernels, runtimes, and model integrations are difficult to build and optimize for every new architecture. Infinity’s investor describes this as a software-tooling bottleneck: alternative chips remain functionally sidelined against NVIDIA’s mature CUDA ecosystem, while hardware teams may spend months or years recreating the systems engineering needed to support modern models. The economic pain is delayed time-to-market, delayed time-to-revenue, and underutilized hardware. The clearest use case is taking a new accelerator—such as d-Matrix’s Corsair—from hardware access to production-quality large-language-model inference quickly.

Product / Service

Infinity provides an automated software layer for AI-chip inference. Its autonomous agent, Ignition, takes a chip’s hardware specifications, reverse-engineers optimization strategies, writes and tests low-level kernels, and generates a complete inference library for the target silicon. The system is designed to work across GPUs, NPUs, TPUs, custom ASICs, and edge accelerators, in whatever programming language the customer’s stack requires, and across text, vision, video, and audio models. Infinity claims a registry of more than 1,000 optimized kernels and roughly 10-times faster time-to-market than manual kernel development.

The delivery model is primarily B2B infrastructure for chip companies, with design partnerships and a cloud deployment that exposes the resulting inference capability. In the d-Matrix engagement, Ignition reached up to 92% of the chip’s empirical compute peak within 10 hours and had Qwen3, Qwen3.5, and Gemma4 running end-to-end within 10 days. The benefit is that hardware vendors can ship competitive, production-ready inference software without building a large specialized kernel-engineering organization from scratch; Infinity’s d-Matrix Cloud is an early live example.

Market

Infinity operates in AI infrastructure, specifically accelerator enablement, inference software, compiler/runtime tooling, and kernel optimization—the software layer between AI silicon and deployed models. Its central competitive reference point is NVIDIA’s CUDA ecosystem, whose advantage comes from decades of hardware-software co-optimization. Other substitutes include chip vendors’ internally built toolchains and general inference stacks such as vLLM; Infinity says its generated engine surpassed vLLM on traditional hardware. Public materials do not identify a single direct startup rival, so the competitive set is best understood as the incumbent CUDA stack, in-house enablement teams, and adjacent compiler and inference frameworks.

Infinity is an early-stage company founded in August 2025. As of July 2026, it had raised $15 million in seed funding at a $100 million post-money valuation, had Touring Capital and Principal VC participation plus strategic angel backing, and reported a live design partnership with d-Matrix alongside active discussions with other major chip companies. Its d-Matrix results and internal-preview cloud provide meaningful technical and commercial validation, but the cited materials do not disclose revenue; it should therefore be viewed as early commercial traction rather than a proven scaled-revenue business.

Founders & Leadership

Jeremy NixonFounder
Founder & CEO
Sravya TirukkovalurHead of Engineering

Funding History

2026-07
Seed$15M

Touring Capital, Principal VC

Recent News

2026-07-20partnership
Partnering to Advance Performant Full-Model Inference

Infinity and d-Matrix partnered to bring Qwen3 from initial tensor-parallel operations to complete, stateful inference on a single d-Matrix Corsair card.

2026-07-20funding
Infinity Raises $15 Million in Seed Funding to Build the Software Layer That Makes Any AI Chip Inference-Ready

Infinity.inc announced a $15 million seed round at a $100 million post-money valuation. The funding will scale its autonomous inference-software agent Ignition, expand engineering, and accelerate chip-company partnerships.

2026-07-20
Inference startup Infinity raises $15M from Touring Capital, OpenAI and Anthropic researchers

TechCrunch reported that Infinity raised $15 million at a $100 million valuation from Touring Capital, Principal VC, and researchers associated with OpenAI and Anthropic. The company is developing chip-agnostic inference software.

2026-07-20
Infinity raises $15M to build a universal AI Inference Library for any Chip

FoundersToday covered Infinity’s $15 million financing at a $100 million valuation and its effort to build a universal inference library for different AI chips.

2026-07-20
Infinity Raises $15 Million Seed Funding At $100 Million Valuation To Build AI Chip Inference Software

Pulse 2.0 reported that Infinity will use the round to scale its automated research platform, expand engineering, and accelerate semiconductor partnerships including its work with d-Matrix. The coverage also highlighted Ignition’s automated kernel-generation and optimization capabilities.

2026-07-20
AI startup Infinity raises $15 million in seed funding at $100 million valuation

Crypto Briefing reported that Infinity raised $15 million to expand its AI-chip optimization platform, with backing from Touring Capital, Principal Venture Partners, chip-industry executives, and OpenAI and Anthropic researchers. Proceeds are intended to grow engineering, improve automated research systems, and deepen d-Matrix partnerships.

2026-04-30product
Igniting d-Matrix: LLM Inference on New Silicon, in Days, Not Years

Infinity described how its Ignition research agent generated a full inference stack for the d-Matrix Corsair, bringing Qwen3, Qwen3.5, and Gemma4 to end-to-end operation on the new processor within 10 days.

Active Roles

1

Business Model

Infinity uses a business-to-business chip-partnership model, providing inference software and enablement to AI-chip designers and hardware partners. The company reports generating millions of dollars in annual recurring revenue from these chip-design partnerships; specific pricing is not disclosed.

Products

Ignition autonomous AI research agentInfinity model-aware inference softwareGenerated production-ready inference libraries for new AI acceleratorsInfy inference optimization systemInfinity d-Matrix Cloud and related hardware-partnership solutions

Customers

d-Matrix (named live design partner; customer status is not explicitly stated)

Tech Stack

Autonomous AI research and optimization agent (Ignition)Generated, hardware-aware AI inference stacksLow-level compute-kernel generation, testing, and optimizationProduction-ready inference libraries and runtimes for AI acceleratorsModel-aware inference optimization for LLMs and multimodal models

Competitors

NVIDIA TensorRT
AMD ROCm
Modular
MulticoreWare

Key Investors

Touring Capital, OpenAI researchers, Anthropic researchers