Companies

Mako

mako.dev

Makora builds hardware-agnostic AI optimization software and inference endpoints that make frontier models faster and cheaper.

HQNew York, New York, United States
Employees1-10
Funding$8.57M
9 active roles
Profile 6mo agoJobs checked 7m ago
AI / MLAI ApplicationB2B SaaSSeed$1M-$10M

About

Mako, now branded Makora, builds AI infrastructure software that automates GPU-kernel generation, inference-engine tuning, algorithmic optimization, and orchestration for developers and enterprise AI teams. Its differentiator is a hardware-agnostic, no-rewrite optimization layer spanning kernels through serving, enabling faster and cheaper frontier-model inference across NVIDIA, AMD, TPUs, Trainium, and other hardware.

Market

Makora competes in AI infrastructure, particularly inference-as-a-service, GPU-kernel generation, and end-to-end inference optimization for open and frontier models. It differentiates by using agents across orchestration, algorithms, serving engines, and GPU kernels—not just hosting models—and by supporting heterogeneous hardware and on-premises deployment, while competitors such as Fireworks, Together AI, Baseten, GroqCloud, and Modal primarily emphasize optimized inference engines, broader AI clouds, or serverless deployment.

Target Customers

Makora serves individual developers and AI engineering teams building agentic coding workflows, as well as enterprise organizations running production AI inference that need dedicated instances, on-premises deployment, custom model serving, or hardware-specific optimization. Its custom-engineering offering also targets model developers, infrastructure teams, and hardware companies that need optimized kernels, inference tuning, RL environments, or embedded design support.

At a Glance

Problem

Mako, also presented in newer materials as Makora, addresses the difficulty and expense of making AI workloads run efficiently on GPUs. GPU kernels—the low-level code that performs the computation—are notoriously difficult to write and optimize, requiring scarce specialist engineering talent. When kernels are inefficient, AI operators waste expensive accelerator capacity, increase inference costs, and sacrifice speed and reliability. The central use case is optimizing kernels for large-scale model inference, particularly for novel model architectures, quantization schemes, or hardware where existing libraries do not deliver good performance.

The economics are substantial because better kernel performance can increase throughput while reducing the amount of GPU time required per request. M13 reports benchmark results claiming a 49% performance improvement and 70% cost reduction for Mistral-7B, and an 85% performance improvement with a 44% cost reduction for Qwen2-72B. These figures are presented as company benchmarks rather than independently verified results, but they illustrate the value proposition: turning low-level optimization work into direct infrastructure savings.

Product / Service

MakoGenerate, spelled MakoraGenerate in the company’s product materials, is an LLM-powered agent that autonomously generates, compiles, validates, and benchmarks GPU kernels. It currently supports CUDA and Triton for NVIDIA GPUs ranging from Hopper to Blackwell. The system combines LLM code generation with an automated feedback loop and evolutionary search: it creates candidate kernels, runs functional tests, collects compiler diagnostics and performance measurements, and feeds that information back into subsequent iterations while exploring choices such as thread-block sizes, memory access patterns, and asynchronous data movement.

The delivery model is a software platform rather than a consulting-only optimization service. The product has been offered as a free research preview and playground, with early access for its more advanced evolutionary-search capabilities; the company’s site also points to inference endpoints that use agents to optimize kernels, inference engines, algorithms, and orchestration. The intended benefit is an abstraction layer in which developers can write or deploy a workload once and obtain hardware-specific implementations without repeatedly hand-writing and tuning kernels across NVIDIA, AMD, or custom accelerators.

Market

Mako competes in AI infrastructure, specifically AI-native compilers, GPU kernel generation, inference optimization, and hardware-agnostic performance middleware. Its stated ambition is to sit between models and hardware, selecting and optimizing the combination of kernel, library, and accelerator for a workload. The closest named comparison in the research is OctoML/OctoAI, an earlier self-optimizing model-deployment platform later acquired by NVIDIA. Luminal is another adjacent competitor: it describes itself as an open-source ML compiler that generates CUDA kernels and simplifies production deployment. Vendor-native stacks such as NVIDIA’s TensorRT-LLM are also adjacent substitutes for teams that prefer a hardware-specific optimization stack, although they are not identical products.

Mako has meaningful financing and technical validation signals but limited publicly disclosed commercial traction in the evidence reviewed. Cornell reports an $8.5 million seed round led by M13, with participation from Neo, Flybridge, and others, while M13 also identifies angel participation including Jeff Dean. The company has published performance benchmarks and made its generator available in research preview, but the available evidence does not provide a customer count, recurring revenue figure, or confirmed production-scale revenue. It is therefore best characterized as an early-stage, funded commercialization effort with a live product and benchmark-led validation, rather than a company with publicly demonstrated scaled revenue.

Founders & Leadership

Waleed AtallahFounder
Co-founder and Chief Executive Officer
Lukasz DudziakFounder
Co-founder and Chief Technology Officer
Mohamed AbdelfattahFounder
Co-founder and Chief Science Officer

Funding History

2024-12
Pre-Seed (inferred)$70K (inferred)

Undisclosed

2025-08
Seed$8.5M

M13

Recent News

2026-07-29product
DSPARK for GLM-5.2: Let’s Draft Longer

Makora describes a custom DSPARK speculative-decoding model for GLM-5.2. Its full training recipe increased accepted length from 3.517 to 3.797 on 50K samples, and a larger run reached 4.915.

2026-07-17product
Reading the Future with Gemma 4

Makora announced an optimized Gemma-4-26B-A4B inference endpoint using a faster attention kernel and its own DSpark drafter. The company reported up to 82% higher throughput on a single AMD MI355X and released the DSpark checkpoint on Hugging Face.

2026-06-29product
One Data Type is Not All You Need for 4-bit Quantization

Makora released the Qwen3.6-35B-A3B-MixFP4 quantized checkpoint and the MixFP4 format, which adaptively selects between NVFP4 and INT4 representations. The checkpoint is about one-third the size of the BF16 model while matching its MMLU-Pro performance, and is available on Hugging Face.

2026-06-05product
Hierarchical SMC-SD: Composing Speculative Decoding Techniques

Makora introduced a hierarchical speculative-decoding approach combining Eagle3 and SMC-SD. In an H100 test, it reached 200.2 tokens per second, a 2.08x speedup over the target model while preserving benchmark accuracy.

2026-05-27product
Maximizing Intelligence per Second: Fast Inference Endpoints for Agentic Systems

Makora announced ultra-fast inference APIs for coding agents and other interactive AI systems. The initial release supports nine open-source models and uses optimization across the GPU-kernel, inference-server, and algorithmic stack.

2025-12-16product
Fast LLM-Generated Kimi Delta Attention Kernels

MakoraGenerate produced functional Kimi Delta Attention kernels using evolutionary search. Makora reported 5.6x–7.8x speedups over Inductor while matching hand-optimized baselines in minutes rather than weeks.

2025-12-03
Mako is now Makora

Mako announced its rebrand to Makora, while retaining the same team and mission of automatically unlocking GPU performance. The company highlighted MakoraGenerate as its platform for automatically writing, optimizing, and deploying GPU code.

2025-08-25
Mako, a faculty-led startup based at Cornell Tech, raises $8.5 million

Cornell reported that the faculty-led startup raised $8.5 million to advance its work optimizing the computing efficiency of graphics and AI workloads.

2025-08-12funding
We Raised $8.5M to Make Peak GPU Performance Universally Accessible

Makora announced an $8.5 million seed round led by M13, with participation from industry leaders. The funding is intended to expand engineering, hardware support, and the company’s GPU optimization platform.

2025-08-12partnership
Makora announces partnerships with AMD and Tenstorrent

Alongside its seed announcement, Makora disclosed partnerships with AMD for kernel generation on ROCm-compatible hardware and with Tenstorrent to improve performance on next-generation silicon.

Active Roles

9
Sydney/Engineering/3d ago
Sydney/Finance/5d ago
Sydney/Finance/21d ago
London/Operations/34d ago
Sydney/Finance/43d ago
Dublin/Finance/70d ago
Dublin/Finance/70d ago
London/Finance/80d ago

Business Model

Makora monetizes its hosted inference platform through token-based pay-as-you-go usage and monthly developer subscriptions, including $20 Starter and $200 Developer plans, with custom-priced enterprise dedicated or on-premises deployments. It also sells custom engineering engagements for kernel optimization, private model serving, reinforcement-learning environments, and embedded design work.

Products

MakoraInference — optimized inference endpoints and model servingMakoraGenerate — AI-agent-based GPU kernel generation and validationMakoraOptimize — automated inference-engine and hyperparameter tuning for vLLM and SGLangCustom Engineering — custom kernels, private model serving, RL environments, and embedded design engineering

Tech Stack

AI agents for GPU optimization and code generationCUDAHIPTriton and other GPU programming DSLsvLLMSGLangTensorRT-LLMGPU kernelsSpeculative decoding, batching, and quantizationGPU orchestration, routing, scheduling, and load balancingNVIDIA, AMD, AWS Trainium, Google TPU, Qualcomm, and Tenstorrent hardware

Competitors

Fireworks AI
Together AI
Baseten
GroqCloud
Modal

Key Investors

M13 Company, Neo, Flybridge