Meta has officially made its SAM 3.1 visual segmentation model available through Model API, giving developers a hosted way to run object segmentation on images and videos without managing model weights, GPUs, or deployment infrastructure themselves. After submitting an image or video and entering a text prompt such as "red bicycle," users can identify the matching object and receive both bounding boxes and pixel-level masks.
The API also supports tracking the same object through video, which Meta says can be used for tasks such as automatic blurring and visual effects. Meta had already released an open-weight version of SAM 3.1 in March this year.
Compared with SAM 3, the biggest change in SAM 3.1 is Object Multiplex. Instead of handling multiple targets one by one, the model can now process up to 16 targets at the same time. In Meta’s tests, a single H100 increased multi-object video processing speed from 16 frames per second to 32 frames per second. When tracking 128 objects, performance was about 7 times faster than SAM 3.
Meta said image segmentation is priced at $2.5 per 1,000 images, while video costs $0.2 per 1,000 frames. The API can also be accessed through the OpenAI SDK, though the hosted version currently accepts only text prompts. The open-weight version still supports visual prompts such as boxes and points.
Meta has officially added its SAM 3.1 visual segmentation model to Model API, giving developers a hosted option for image and video segmentation.
After submitting an image or video and entering a text prompt such as "red bicycle," developers can identify the matching object and receive bounding boxes and pixel-level masks.
For video, the API can keep tracking the same target across frames. Meta said this can help with tasks such as automatic blurring and adding visual effects.
Meta had already released an open-weight version of SAM 3.1 in March this year. Compared with SAM 3, the biggest upgrade in SAM 3.1 is Object Multiplex. Multiple targets no longer need to be processed one at a time. The system can now handle as many as 16 targets in a single run.
In Meta’s tests, one H100 increased multi-object video processing speed from 16 frames per second to 32 frames per second. When tracking 128 objects, speed was about 7 times that of SAM 3.
Previously, developers had to download weights themselves and prepare GPU resources and deployment environments. With the new setup, they can call Meta’s hosted API directly. Pricing is set at $2.5 per 1,000 images for image segmentation and $0.2 per 1,000 frames for video. The API can also be accessed through the OpenAI SDK.
The API version currently accepts only text prompts. The open-weight version still supports visual prompts such as boxes and points for target selection. The video API can track up to 16 objects per frame.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.