Companies

RightNow

runinfra.ai

RightNow builds GPU-native infrastructure that optimizes open AI models and deploys them efficiently for teams.

HQMiddletown, Delaware, United States
Employees1-50
Jobs checked 16h ago
Cloud InfrastructureAI InfrastructureInfrastructure

About

RightNow is a research lab building GPU-native infrastructure that helps AI teams own and operate their model stack rather than depend on closed-source APIs. Its RunInfra platform accepts open models, optimizes GPU kernels and runtime settings, benchmarks hardware, and deploys production inference through managed or self-hosted infrastructure.

Market

RightNow competes in AI inference infrastructure: the market for optimizing, serving, and operating open-source models on GPUs as production APIs. RunInfra differentiates itself from managed inference providers such as Baseten, Together AI, Modal, and Fireworks AI by combining chat-native model and GPU selection, measured kernel/runtime optimization, and deployment in one workflow. It also offers an exportable self-hosted stack with no vendor lock-in, alongside managed scale-to-zero infrastructure and pay-per-token pricing.

Target Customers

Primary customers are AI/ML product teams—especially startups and enterprise software or industry teams—whose ML engineers and platform/infrastructure leads need to turn open-source models into production APIs. The self-serve Core plan fits smaller teams, while Enterprise targets larger organizations requiring dedicated GPU infrastructure, self-hosted or custom-GPU deployment, RBAC, compliance, and service-level agreements.

At a Glance

Problem

RightNow addresses the gap between open AI models and the GPU infrastructure required to run them efficiently in production. Teams must choose compatible models and GPU tiers, generate or tune kernels, configure inference runtimes, and balance latency against cost; that work is specialized, time-consuming, and can leave companies dependent on closed-source model APIs or generic hosting. The central use case is taking an open-source or Hugging Face model and turning it into a production-ready API optimized for the customer's workload and hardware.

The economic pain is variable inference cost and performance: the right combination of quantization, speculative decoding, KV-cache reuse, FlashAttention, batching, and GPU selection can materially affect throughput, latency, and spend. RunInfra frames the problem as finding the cheapest GPU that meets a latency target, rather than simply renting standard infrastructure.

Product / Service

RunInfra is RightNow's chat-native model-optimization and inference-infrastructure platform. A user describes the model or AI application they want to run; RunInfra selects compatible open models, benchmarks GPU tiers, generates optimized GPU kernels, tunes supported runtime settings, and produces a deployment-ready API. Its optimization surface includes quantization, speculative decoding, KV-cache reuse, FlashAttention, continuous batching, and server tuning, with performance measured on the selected GPU.

The service can be consumed as managed infrastructure through scale-to-zero, OpenAI-compatible API endpoints and pay-per-token serving, or exported for self-hosting on a customer's cloud or bare-metal GPUs. This gives teams a faster path from model selection to production while preserving more control than a closed API: the stack is inspectable and exportable, and enterprise plans add dedicated infrastructure, custom GPU deployment, compliance controls, and SLAs.

Market

RightNow competes in the AI inference infrastructure and model-optimization market, alongside managed open-model platforms such as Baseten, Fireworks AI, and Together AI. Those companies also serve and scale open or custom models, while RunInfra differentiates around model-hardware co-design: automatically generating kernels, benchmarking the actual GPU path, and offering an exportable runtime rather than only a hosted model endpoint. Its self-serve pricing starts with monthly credits, while enterprise customers receive custom infrastructure and commercial terms.

The product appears commercially launched but early. RunInfra was listed on Product Hunt on June 30, 2026, with claims of faster and cheaper execution than standard hosting; the official site offers free onboarding, paid Core plans, and enterprise plans. Y Combinator lists RightNow as an active company founded in 2025 with a two-person team in Amman and a Fall 2026 batch affiliation. The gathered evidence does not report revenue, customer names, or usage metrics, so it is best characterized as an early commercial product with public pricing and launch traction rather than a company with publicly demonstrated scale.

Founders & Leadership

Jaber JaberFounder
Technical Co-founder
Osama JaberFounder
Commercial Co-founder

Funding History

2026-08
Seed$8M

Y Combinator

Recent News

2026-07-31
$0.09 and $290.12: What Actually Moves Your Inference Bill

RunInfra compares $0.09 and $290.12 as prices for one million output tokens, highlighting the wide variation in inference costs for the same billing unit.

2026-06-30product
RunInfra: Describe the AI model you need and get an optimized AI

RunInfra launched on Product Hunt as a chat-native infrastructure tool that turns plain-language descriptions of open-source models or applications into production APIs. It benchmarks GPUs, quantizes models, and generates custom CUDA kernels, with pay-per-million-token pricing and scale-to-zero deployment.

2026-06-21
StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k

RunInfra featured its StreamIndex research on memory-bounded sparse attention, which selects top-k keys in a streaming pass and uses a fused Triton kernel for production inference workloads.

2026-06-20product
Deploy your first optimized model, measured before you ship

RunInfra published a product or engineering update describing a workflow in which users describe a goal and RunInfra builds and optimizes the inference stack before deployment.

2026-05-22
RightNow: Enabling Model-Hardware Co-Design at Scale

Y Combinator’s company profile describes RightNow AI as a research lab building GPU infrastructure and says RunInfra accepts Hugging Face models, generates optimized GPU kernels, and deploys them serverlessly with pay-per-token pricing.

2025-09-03
RightNow AI presents its CUDA editor, RunInfra infrastructure, and Forge kernel work

RightNow’s company site identifies it as a YC-backed GPU research lab and highlights its CUDA editor, RunInfra inference infrastructure, Forge kernel work, and AutoMegaKernel research.

Active Roles

0

No active roles right now.

Get notified when they post

Business Model

RunInfra monetizes through monthly credit subscriptions covering optimization, deployments, and agent usage, plus usage-based managed GPU serving priced per million tokens. It also offers enterprise plans with custom pricing, infrastructure, credit volumes, compliance, and support terms.

Products

RunInfra chat-native model optimization and inference platformRunInfra Cloud: managed, autoscaling GPU deployment with pay-per-token pricingRunInfra Self-Hosted: exportable optimized kernels and runtime for customer-owned cloud, bare-metal, or on-premises GPUsRightNow Editor: an AI-powered code editor for NVIDIA GPU hardware development

Customers

NVIDIAAMDMITTogetherGoogleRunway

Tech Stack

Open-source foundation models from Hugging Face, spanning text, image, speech, vision, embeddings, and audiovLLM, SGLang, TensorRT-LLM, and vLLM-Omni inference runtimesGPU kernel generation and optimization, including FlashAttention v2, PagedAttention, speculative decoding, and KV-cache reuseAWQ, GPTQ, and FP8 quantizationNVIDIA GPU infrastructure, including T4, L4, L40S, A100, H100, H200, and B200OpenAI-compatible APIs with autoscaling and scale-to-zero deployment

Competitors

Baseten
Together AI
Modal
Fireworks AI
RunPod