Banbury Road unveils 32-model Kardashev-0.7 trained with RLPS

Banbury Road unveils 32-model Kardashev-0.7 trained with RLPS

N
News Editor
2026-10-06 15:18:16
AI lab Banbury Road has introduced Kardashev-0.7, a system made up of 32 different models trained together with a method it calls RL for Population Scaling, or RLPS. Instead of assigning fixed roles in advance, the team said it used reinforcement learning to let the models develop distinct specialties and complementary capabilities during training. The company compared the approach with earlier RLPS experiments. In an 8-model setup, the jointly trained population scored 81.70%, above 71.04% for eight independently trained models and 72.65% for sampling the same model eight times. In a 16-model experiment, 22 questions were answered correctly by only one model in the group, which Banbury Road said suggests the models learned different capabilities. Kardashev-0.7 expands the scale to 32 models. Banbury Road said the system can deliver frontier-model-level performance at 0.007x to 0.02x inference cost and 0.03x memory use. Its API beta is still gated behind a waitlist. The company has not yet disclosed a full method for how the system would automatically select the correct answer in real-world use, since earlier group scores were produced by choosing the best output after the fact using standard answers.

Banbury Road has released Kardashev-0.7, a system built from 32 different models. The lab said it trained the models together with reinforcement learning, without assigning fixed responsibilities in advance, allowing them to develop distinct specialties and complementary strengths during training.

The approach is called RL for Population Scaling, or RLPS. Earlier RLPS experiments had reached 16 models at most. Kardashev-0.7 pushes that scale to 32.

Results from 8-model and 16-model experiments

In an 8-model experiment, the jointly trained model population scored 81.70%. That was higher than 71.04% for eight independently trained models and also above 72.65% for sampling the same model eight times.

In a 16-model experiment, 22 questions were answered correctly by only one model in the group. Banbury Road said this indicates the models did in fact learn different capabilities.

Cost claims and API access

Banbury Road said Kardashev-0.7 can reach frontier-model-level performance at 0.007x to 0.02x inference cost and 0.03x memory use. The API beta currently requires users to join a waitlist.

An unresolved deployment question

Banbury Road also noted an important limitation. In the earlier 8-model and 16-model experiments, the group scores were calculated after the fact by selecting the best answer from multiple model outputs using standard answers. In actual use, there is no standard answer available, and the company has not disclosed a full method for how the system would automatically choose the correct response.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.