Japanese developer Kamimoto has integrated Jev into the MiniMax H3 video generation pipeline, introducing a dynamic way to control how much Attention is kept at different layers. Instead of using a fixed ratio, the system selects from 1%, 3%, 5%, and 10% based on each layer’s runtime state, while Sparse Attention determines which calculations can be skipped. In a test on an RTX 4070 12GB, generating a roughly five-second video with the original four-step setup took 6 minutes and 7 seconds. After adding Sparse Attention and Jev-based dynamic control, the time dropped to 3 minutes and 34 seconds, a 41.7% reduction. The report also makes clear that the full gain should not be attributed to Jev alone, because Sparse Attention already speeds up generation by reducing part of the Attention workload, and fixed 5% or 10% retention settings also improve speed. At this stage, there is still no proof that dynamic adjustment delivers better image quality than a fixed ratio. The author said quality has not yet been fully validated, and the experiment code has been open-sourced.
Japanese developer Kamimoto has connected Jev to the MiniMax H3 video generation workflow. Under this setup, Jev chooses how much Attention to retain at each layer based on runtime conditions, selecting from 1%, 3%, 5%, and 10%. Sparse Attention then decides which specific calculations can be skipped.
Benchmark result
On an RTX 4070 12GB, generating a video of about five seconds took 6 minutes and 7 seconds with the original four-step version. After adding Sparse Attention and Jev-based dynamic control, the time fell to 3 minutes and 34 seconds, reducing total generation time by 41.7%.
What drove the speedup
The report notes that the full 41.7% improvement should not be credited to Jev alone. Sparse Attention already speeds up the process by cutting part of the Attention computation, and fixed retention levels such as 5% or 10% also make generation faster. Jev’s role is to adjust the ratio dynamically, allocating more computation to important layers and less to less important ones.
For now, there is no proof that this dynamic method produces better visual quality than a fixed ratio. The author also said image quality has not been fully validated. The experiment code has already been released as open source.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.