Taiwanese YouTuber Joeman said in a YouTube video uploaded on Oct. 1 that he bought a fully configured M5 Ultra Mac Studio for NT$437,900. He said his team tested the machine for a week across benchmarks, video editing, local AI workloads, and gaming. The 18-minute video had drawn about 560,000 views by the afternoon of Oct. 3.
His takeaway was direct: the machine’s AI performance was unusually strong. Joeman said local large language model inference and image generation were already close to cloud-model speed, while video generation still trailed top cloud services by a visible margin.
A fully loaded M5 Ultra configuration priced at NT$437,900
Joeman chose the highest-end M5 Ultra variant with a 36-core CPU and 80-core GPU, then upgraded unified memory to 256GB and storage to 4TB. On Apple Taiwan’s website, the standard M5 Ultra version, with a 30-core CPU, 64-core GPU, 96GB of memory, and 1TB of storage, starts at NT$199,900.
The higher-spec chip version starts at NT$245,400. Joeman said the memory upgrade added NT$140,000 and the storage upgrade added NT$52,500, bringing the total to NT$437,900. For comparison, the cheapest Mac Studio model with an M5 Max starts at NT$84,900, making his unit roughly a five-times pricier top-end configuration.
He also pointed to Apple silicon’s unified memory design, where the CPU and GPU share the same memory pool. Joeman said the first bottleneck for large AI models on a typical PC is often graphics memory. He added that the top consumer-grade RTX 5090 has 32GB, while a professional workstation GPU can reach 96GB but costs about NT$590,000, more than this Mac Studio. He also said NVIDIA still holds an advantage in AI compute and model training, while the Mac Studio stands out for memory capacity and lower power draw.
Editing export time was cut by more than half versus M4 Max
The benchmark and editing comparison used Joeman Studio’s previous top machine, an M4 Max MacBook Pro. In Geekbench, the M5 Ultra posted 3,767 in single-core and 52,569 in multi-core. Joeman said that was 15% higher in single-core performance and close to double in multi-core performance compared with the M4 Max.
In Cinebench 2024, multi-core performance was 183% higher, or about 2.8 times the M4 Max. In 3DMark graphics testing, scores came in at 2.5 to 3 times the M4 Max.
For editing, the team used Premiere with a 20-track 4K project. Joeman said timeline scrubbing remained smooth. Exporting the project to a high-bitrate H.264 4K file took 3 minutes 2 seconds on the M5 Ultra, versus 6 minutes 27 seconds on the M4 Max. Final Cut exported 4K footage 2 minutes 16 seconds faster, and Blender rendering finished 4 minutes 50 seconds faster.
Gaming performance landed between the RTX 5060 and RTX 5070
On gaming, Joeman tested the native Mac version of Cyberpunk 2077. With ray tracing set to the highest level, average frame rate came to 46 FPS at 2K resolution and 24 FPS at 4K. With Apple’s MetalFX upscaling enabled, performance rose to 78 FPS at 2K and 57 FPS at 4K.
Black Myth: Wukong, run through CrossOver translation, reached 69 FPS at 2K on high settings. Joeman’s conclusion was that the M5 Ultra’s gaming performance sat between the RTX 5060 and RTX 5070.
Local text and image generation neared cloud processing speed
For local AI tests, Joeman compared the system with M2 Max and M3 Max MacBook Pro models using more mainstream memory configurations. He said a top-end M4 Max costs about NT$240,000, putting it too far out of reach for most users. He also explained that the “B” in model names stands for billions of parameters.
Using DeepSeek as an example, Joeman said a MacBook with 64GB of memory could only run a compressed 70B version, while the 256GB M5 Ultra could load a 671B version. In his view, the first difference between the two was simply whether the model could run at all. Follow-up tests used models at the same precision across all three machines.
In matrix compute tests measured in TFLOPS, the M2 Max and M3 Max scored between 9 and 11, while the M5 Ultra reached 94 to 95. Joeman cautioned that higher compute numbers do not mean daily workloads become 10 times faster, so he moved on to practical tasks.
Text generation used a 30B MoE model, with MoE referring to a mixture-of-experts architecture that activates only part of the parameters for each answer. Under a heavy 30K-input workload, the waiting time after sending a prompt was 2 minutes on the M2 Max, about 1 minute on the M3 Max, and just 8.6 seconds on the M5 Ultra, roughly one-seventh of the M3 Max. Once generation began, the M5 Ultra produced about 80 tokens per second, around 2.5 times the M3 Max. Joeman noted that tokens are the basic units AI uses to process text and can be a character or part of a word.
In a coding test, response times for a date-filtering task were about 7.2 seconds on the M2 Max, 4.7 seconds on the M3 Max, and 3.4 seconds on the M5 Ultra.
For image generation, the team used a Qwen-series image model to turn a photo of Joeman into a sci-fi spaceship scene. The M2 Max took 3 minutes 4 seconds, the M3 Max took 49 seconds, and the M5 Ultra finished in 27 seconds. Joeman said that speed was already on par with cloud processing.
Local video generation improved sharply but still trailed leading cloud models
Video generation produced the biggest spread among the three machines. Using Wan 2.2 to create a 1.7-second 480p video from text, the M2 Max took 8 minutes, the M3 Max 4 and a half minutes, and the M5 Ultra 38 seconds. At 704p, the M2 Max took 17 minutes while the M5 Ultra took 1 minute 40 seconds.
Against cloud services, though, the M5 Ultra remained slower in video generation. Creating a 5-second clip took 6 minutes locally on the M5 Ultra. Cloud Wan 2.2 took 5 minutes 7 seconds, Kling 3.0 took 3 minutes 52 seconds, and two other cloud models came in at 2 minutes 26 seconds and 3 minutes 56 seconds.
Joeman said the M5 Ultra was close to the cloud version of Wan 2.2, but still well behind top-tier cloud video models. He added that actual cloud speed also depends on server load at the time.
He also estimated costs using official pricing from each provider and an exchange rate of US$1 to NT$32. A 5-second video on Kling 3.0 came to about NT$14, while other cloud models ranged from NT$14 to NT$55. Local generation only required electricity. At roughly NT$55 per cloud-generated clip, he calculated that the price of the computer would equal about 8,000 five-second videos, with NT$437,900 divided by NT$55 coming to 7,962 clips. Joeman also said local video quality still falls short of the best cloud models, though the advantage is that users can keep iterating until they get what they want.
Memory usage nearly filled the machine, and Joeman expects local and cloud AI to coexist
In a memory stress test, Joeman ran three Qwen-series models at the same time for text, coding, and image generation. Memory usage peaked at 245GB, nearly filling the machine. He said work that once needed three computers could now be handled on one system.
He also ran a 671B quantized DeepSeek model. Quantization reduces memory use by compressing model parameters to lower precision. In a task that read fictional warehouse shipment data and checked quantities, estimated peak memory use reached 240GB, or about 94% of total system capacity.
Joeman said the machine had more than enough headroom for video editing, Blender, and model workloads. At the same time, he said an M5 Max, or even the previous-generation M3 Ultra, would probably already be enough for many of those jobs. In his view, the main story of this generation is the jump in AI performance.
He said that before testing, neither he nor his editing team believed a local machine could match cloud models backed by far larger compute resources. After running the tests, he came away with a different view: language model inference and text-to-image generation were already very close to cloud-level speed. He estimated that local image quality reached about 90% to 95% of the detail seen in top cloud models, making it practical to generate repeatedly on-device, choose the preferred result, and then use cloud tools for final touch-ups. Video generation, he said, still has a clear gap because the compute load is much heavier.
Joeman argued that the biggest value of the Mac Studio is confidentiality, since all data stays on the machine. He said companies sometimes sign nondisclosure agreements when unboxing products, and early product units that need benchmarking or analysis are not suitable for cloud-connected workflows. For small and medium-sized businesses handling sensitive data, he said running as much as possible locally is the better option.
His broader view is that local and cloud computing will split the workload and coexist. He expects local compute to keep improving, with some tasks handled on-device and others left to top-end cloud infrastructure.

