RoCE

Meta AI
2026-08-25 17:27:21

Meta AI unveils MetaRoCE, an RDMA transport built for AI-scale Ethernet

Meta AI has introduced MetaRoCE, a new RDMA transport protocol designed for AI workloads running over commodity Ethernet. According to Techub, the protocol departs from standard RoCE by treating the network as lossy rather than lossless, while shifting packet ordering, path selection, and recovery functions to the NIC layer. The stated goal is to address network bottlenecks that emerge during training across large-scale AI clusters. Meta said it plans to release the specification, a reference software implementation, and a compliance test suite through the Open Compute Project. Hardware support remains at an early stage. The protocol has already been validated on AMD Pensando programmable NICs, and implementations from other vendors are still in progress. The report said the announcement should be viewed more as an architectural decision than a purchasing one. MetaRoCE’s design choices include out-of-order delivery by default, native multipathing, fault tolerance in place of lossless networking, end-to-end congestion control, and topology independence. Techub cited MarkTechPost for the report, which said the protocol builds directly on Meta’s 2024 work on large-scale RoCE and its broader infrastructure evolution.

200
Meta AI unveils MetaRoCE, an RDMA transport built for AI-scale Ethernet
AMD
2026-08-04 11:15:08

Wafer says Kimi K3 fits on one 8x AMD MI355X server, while B200 needs 16 GPUs across two nodes

Wafer AI said it deployed Kimi K3 on AMD’s MI355X, claiming the model can run on a single server with eight GPUs, while the same workload requires 16 NVIDIA B200 GPUs spread across two servers because of memory limits. In Wafer’s test with 1,024 input tokens and 400 output tokens, the MI355X setup delivered 952 tokens per second of aggregate throughput and 118 tokens per second for single-user generation. The article says that, on a single-node basis, this works out to roughly 3.8 times the throughput of the 16-card B200 deployment’s average per-node figure, while B300 still leads on absolute speed at 1,568 tokens per second and 172 tokens per second for single-stream generation. Wafer also published a cost comparison using hourly prices of $2.5 for MI355X, $4.25 for B200, and $6 for B300, under which MI355X offered about 48 tokens per second per dollar, ahead of B200’s roughly 7 and B300’s roughly 33. The report also highlights ROCm software progress, saying Kimi K3 was largely able to run directly on MI355X, with follow-up work focused on a small number of compatibility issues and performance tuning, including speculative decoding fixes and a prefill optimization that reduced cold-start latency on a 172,000-token task.

2260
Wafer says Kimi K3 fits on one 8x AMD MI355X server, while B200 needs 16 GPUs across two nodes