About
RightNow is a research lab building GPU-native infrastructure that helps AI teams own and operate their model stack rather than depend on closed-source APIs. Its RunInfra platform accepts open models, optimizes GPU kernels and runtime settings, benchmarks hardware, and deploys production inference through managed or self-hosted infrastructure.
Market
RightNow competes in AI inference infrastructure: the market for optimizing, serving, and operating open-source models on GPUs as production APIs. RunInfra differentiates itself from managed inference providers such as Baseten, Together AI, Modal, and Fireworks AI by combining chat-native model and GPU selection, measured kernel/runtime optimization, and deployment in one workflow. It also offers an exportable self-hosted stack with no vendor lock-in, alongside managed scale-to-zero infrastructure and pay-per-token pricing.
Primary customers are AI/ML product teams—especially startups and enterprise software or industry teams—whose ML engineers and platform/infrastructure leads need to turn open-source models into production APIs. The self-serve Core plan fits smaller teams, while Enterprise targets larger organizations requiring dedicated GPU infrastructure, self-hosted or custom-GPU deployment, RBAC, compliance, and service-level agreements.
At a Glance
Problem
RightNow addresses the gap between open AI models and the GPU infrastructure required to run them efficiently in production. Teams must choose compatible models and GPU tiers, generate or tune kernels, configure inference runtimes, and balance latency against cost; that work is specialized, time-consuming, and can leave companies dependent on closed-source model APIs or generic hosting. The central use case is taking an open-source or Hugging Face model and turning it into a production-ready API optimized for the customer's workload and hardware.
The economic pain is variable inference cost and performance: the right combination of quantization, speculative decoding, KV-cache reuse, FlashAttention, batching, and GPU selection can materially affect throughput, latency, and spend. RunInfra frames the problem as finding the cheapest GPU that meets a latency target, rather than simply renting standard infrastructure.
Product / Service
RunInfra is RightNow's chat-native model-optimization and inference-infrastructure platform. A user describes the model or AI application they want to run; RunInfra selects compatible open models, benchmarks GPU tiers, generates optimized GPU kernels, tunes supported runtime settings, and produces a deployment-ready API. Its optimization surface includes quantization, speculative decoding, KV-cache reuse, FlashAttention, continuous batching, and server tuning, with performance measured on the selected GPU.
The service can be consumed as managed infrastructure through scale-to-zero, OpenAI-compatible API endpoints and pay-per-token serving, or exported for self-hosting on a customer's cloud or bare-metal GPUs. This gives teams a faster path from model selection to production while preserving more control than a closed API: the stack is inspectable and exportable, and enterprise plans add dedicated infrastructure, custom GPU deployment, compliance controls, and SLAs.
Market
RightNow competes in the AI inference infrastructure and model-optimization market, alongside managed open-model platforms such as Baseten, Fireworks AI, and Together AI. Those companies also serve and scale open or custom models, while RunInfra differentiates around model-hardware co-design: automatically generating kernels, benchmarking the actual GPU path, and offering an exportable runtime rather than only a hosted model endpoint. Its self-serve pricing starts with monthly credits, while enterprise customers receive custom infrastructure and commercial terms.
The product appears commercially launched but early. RunInfra was listed on Product Hunt on June 30, 2026, with claims of faster and cheaper execution than standard hosting; the official site offers free onboarding, paid Core plans, and enterprise plans. Y Combinator lists RightNow as an active company founded in 2025 with a two-person team in Amman and a Fall 2026 batch affiliation. The gathered evidence does not report revenue, customer names, or usage metrics, so it is best characterized as an early commercial product with public pricing and launch traction rather than a company with publicly demonstrated scale.
Founders & Leadership
Funding History
Y Combinator
Recent News
RunInfra compares $0.09 and $290.12 as prices for one million output tokens, highlighting the wide variation in inference costs for the same billing unit.
RunInfra launched on Product Hunt as a chat-native infrastructure tool that turns plain-language descriptions of open-source models or applications into production APIs. It benchmarks GPUs, quantizes models, and generates custom CUDA kernels, with pay-per-million-token pricing and scale-to-zero deployment.
RunInfra featured its StreamIndex research on memory-bounded sparse attention, which selects top-k keys in a streaming pass and uses a fused Triton kernel for production inference workloads.
RunInfra published a product or engineering update describing a workflow in which users describe a goal and RunInfra builds and optimizes the inference stack before deployment.
Y Combinator’s company profile describes RightNow AI as a research lab building GPU infrastructure and says RunInfra accepts Hugging Face models, generates optimized GPU kernels, and deploys them serverlessly with pay-per-token pricing.
RightNow’s company site identifies it as a YC-backed GPU research lab and highlights its CUDA editor, RunInfra inference infrastructure, Forge kernel work, and AutoMegaKernel research.
Active Roles
0No active roles right now.
Get notified when they postBusiness Model
RunInfra monetizes through monthly credit subscriptions covering optimization, deployments, and agent usage, plus usage-based managed GPU serving priced per million tokens. It also offers enterprise plans with custom pricing, infrastructure, credit volumes, compliance, and support terms.