Wafer, an AI inference startup with just eight employees, has raised $40 million in a Series A round, valuing the company at over $200 million. The company turned down acquisition offers from several cloud providers and inference service providers before the round.
Wafer's inference business reached $8 million in annualized recurring revenue (ARR) within about three months of launch. The company does not manufacture chips; instead, it helps models run faster and cheaper on existing hardware.
AI-Driven Optimization
Wafer's core technology uses an AI agent to automatically adjust models, inference engines, kernels, caching, quantization, and scheduling, and to find the optimal configuration for different hardware such as NVIDIA and AMD. Previously, much of this inference performance optimization was done manually by engineers. Wafer aims to automate this process.
In internal tests, running GLM 5.2 on AMD MI355X, Wafer achieved about 80% of the throughput of NVIDIA B200, with less than half the cost.
Customers and Market Validation
Current customers include Vercel and Inworld AI. Vercel independently published test results showing that Wafer's throughput for GLM 5.2 was about twice that of other serverless providers.
Expansion After Funding
After the funding, Wafer immediately began hiring. It opened four roles: engineering, growth, CEO office, and GTM. All positions require working five days a week in office in San Francisco. The engineering role offers a base salary of $250,000 plus equity.

