v1.7.50
Apple system frameworks remain under the central commercial hold. This release makes chat and vision jobs start sooner and stream as they generate.
- Chat replies now stream to the customer as your Mac generates them, instead of arriving all at once when the job ends.
- After a chat job, your Mac keeps that one model loaded for up to 10 idle minutes, so the next request for it skips reloading and re-checking the weights. Everything from the request is still wiped after every job, any other kind of job unloads the model first, and the model is released at once if macOS reports memory pressure.
- Chat and vision jobs start sooner: a model already on your Mac is checked in full once, inside the isolated worker, instead of also being re-read before every job, and the signed model authorization is reused for up to two minutes instead of being downloaded twice per job.
- Your Mac can run 8-bit versions of popular models (Qwen3, Llama 3.3 70B, GLM-4.5 Air) and, with 192 GB of memory or more, large models: Qwen3 235B, GLM-4.6, Qwen3 Coder 480B and DeepSeek V3.1.
- Each job reports how long each step took on your Mac (model check, load, first token, generation), with no prompt or result content, so slow steps can be found and fixed.
- When few Macs hold a released model, the network may ask an idle Mac that can run it to download it in the background, so the next request for it does not wait on a download. Your Mac does this only while it is idle and ready for work, never for a model you turned off auto-download for on the Models page, one model at a time, and only after the usual disk and memory checks.