Spanish AI startup Speridlabs has tested an image-generation approach that skips the compressed latent stage and generates every pixel directly. The team wanted to see whether removing that intermediate step would preserve more detail and improve results on tasks that depend on fine visual information. To run the experiment, Speridlabs trained an open-source model called Iris-3B and also modified another existing model into a direct-pixel version.
The models were then evaluated on tasks including low-resolution image restoration and judging the relative distance of objects in photos, with results compared against conventional methods. The outcome did not show the stable improvement the team had expected. Some metrics came in slightly worse, and the direct-pixel models produced mild checkerboard artifacts in parts of the testing.
Even so, the work showed that direct pixel generation is technically feasible. In one image-generation benchmark, Iris-3B scored 0.540, essentially level with Qwen-Image at 0.539. The findings suggest that removing image compression alone did not give the model a clear edge in tasks that demand stronger detail fidelity.
Spanish AI startup Speridlabs has tried a different route for image generation: skipping the compressed image representation used by many models and generating each pixel directly.
Many current image-generation systems first create images in compressed visual information and then reconstruct them into full pictures. Speridlabs wanted to test whether removing that step would improve output quality by avoiding detail loss.
To do that, the team trained an open-source model called Iris-3B and also converted another existing model into a version that generates pixels directly. The models were tested on low-resolution image restoration and on judging how near or far objects appear in photos, then compared with the original approach.
The results did not show the stable improvement the team had expected. Some metrics were slightly worse. In some tests, the direct-pixel models also showed mild checkerboard patterns.
The study still showed that direct pixel generation is feasible. Iris-3B scored 0.540 on one image-generation benchmark, nearly matching Qwen-Image at 0.539. On the whole, removing the image compression step by itself did not give the model a clear advantage in tasks that require finer detail.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.