FLock2026-09-13 09:13:39FLock Research gets three papers accepted to ECML PKDD 2026 on complex LLM reasoningFLock said three papers from its research team have been accepted to ECML PKDD 2026, a major international conference in machine learning, data mining, and knowledge discovery in Europe. The company said the papers focus on recurring problems in complex reasoning by large models, including shortcut-taking, reasoning errors, and drift away from the original objective during long multi-step chains of thought. According to FLock, one of the studies helps models stay closer to their original line of reasoning during extended multi-step inference. It reduced reasoning drift by 63% and improved average accuracy by 7.4 percentage points across models of different sizes. The other two papers target reasoning on unfamiliar problems and the efficiency of partial correction after a model detects an error in its reasoning process. FLock Research said it continues to conduct foundational work around large-model training, inference, and reliability, with the stated goal of improving accuracy, stability, and generalization on complex tasks so models can move from answering questions to reliably solving more complex problems.820
OpenRouter2026-09-10 01:51:30OpenRouter lists Mistral Small 4 (batch) modelOpenRouter has added a new model, "Mistral: Mistral Small 4 (batch)," according to a Techub News item citing an official OpenRouter announcement. The model is described as the next major version in the Mistral Small series. OpenRouter said it brings together the capabilities of multiple flagship Mistral models in a single system and is built with strong reasoning performance in mind. The announcement also noted that models often appear on OpenRouter before vendors make a formal public announcement. No further technical details were provided in the brief. The update centers on the model’s availability on the OpenRouter platform and the positioning of Mistral Small 4 within the broader Mistral Small lineup.200
Meta2026-09-05 15:35:44Meta Officially Opens Muse Spark 1.3 Max, the Strongest Reasoning TierMeta has released Muse Spark 1.3 Max, the highest reasoning tier of the model. Previously limited to partners due to safety testing, it is now available in Muse Code and the Meta Model API. Chief AI officer Alexandr Wang says it outperforms the high and xhigh tiers on coding and agent tasks. The model scored 68 on the Coding Agent Index, entering the top tier.770
Google2026-08-16 16:02:49Gemini 3.7 Flash review: big coding gains, weaker reasoning and writing still showGoogle launched Gemini 3.7 Flash on August 13 and made it generally available in more than 160 countries on day one. According to Decrypt’s review, the model accepts up to 1 million input tokens, returns 64,000 output tokens, handles images, video, audio, and PDFs, and can use tools while operating a computer. Google’s own benchmark sheet says the model beats Claude Sonnet 5 and GPT-5.6 Terra in 11 of 18 tested categories, including 1,588 Elo on Code Arena’s web development board and 30.4% on AutomationBench, though Decrypt notes those figures come from Google’s methodology and should be treated as company claims rather than settled fact. Decrypt’s hands-on tests found the sharpest improvement in coding. Gemini 3.7 Flash generated a playable browser game on the first try in 2 minutes and 13 seconds, a major step up from Gemini 3.6 Flash, which Decrypt said could not produce a working file in a similar test after its July 21 release. Results were less convincing elsewhere. In creative writing, Decrypt said Gemini produced a tidy story but broke the central prompt rule, losing to a free community model, Qwopus3.5-27B-v3. In associative reasoning, logic, and advanced math, the review said Gemini often showed decent structure but failed on crucial task requirements, including a bridge puzzle and a polynomial problem it left unfinished. Decrypt’s conclusion: Gemini 3.7 Flash is a strong low-cost execution model inside Google’s ecosystem, but its creativity and reasoning remain uneven.1560