Chat and LLM inference
Stream chat completions through an OpenAI-compatible endpoint. Model availability is checked when you run a request.
Open chat →OpenAI-compatible chat, vision Q&A, and fine-tuning. Requests are matched to a compatible provider device and quoted before they run. Live capacity is checked at request time.
Unified memory lets the CPU and GPU work from the same pool of data…
Keep the chat client you already use.
See the unit rate and estimated total, then set a hard spend cap.
MLX workloads on compatible Mac computers with Apple silicon.
Discover workload, model, and lifecycle status.
Chat is the front door. The same account and API also handle vision Q&A and fine-tuning as your application grows.
Stream chat completions through an OpenAI-compatible endpoint. Model availability is checked when you run a request.
Open chat →Ask questions about an image with a model selected from the live catalog.
Open vision Q&A →Create lightweight adapters for supported language models.
Open fine-tuning →Upscale an image to four times its width and height. You get a PNG back.
Open upscale →Turn speech into timestamped text in 99+ languages, with the language detected for you.
Open transcription →Generate SRT or WebVTT subtitles timed to the speech in an audio or video file.
Open subtitles →Change the base URL, choose an available model, and stream the response. The same account gives you the native SDK, CLI, MCP server, and generic job API.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["CC_API_KEY"],
base_url="https://api.commoncompute.ai/v1",
)
stream = client.chat.completions.create(
model="qwen3-4b",
messages=[{"role": "user", "content": "Explain unified memory."}],
extra_body={
"data_class": "public",
"marketplace_execution_risk_acknowledged": True,
},
stream=True,
)
for event in stream:
print(event.choices[0].delta.content or "", end="")Common Compute routes each workload to a Mac computer with the memory, operating system, and runtime it needs. The currently offered catalog uses MLX workloads and explicitly matched model requirements.
Language and generative models run through an Apple Silicon-native machine learning stack.
Choose from the models currently listed for chat, vision, and fine-tuning workloads.
The router checks workload, model, memory, and runtime requirements before dispatch.
Jobs run on independently operated, provider-owned Mac computers. Common Compute records the selected device, price, usage, and receipt evidence so you can review what happened after a job completes.
Read the trust modelBefore: see a quote and choose a spend limit.
During: the router matches explicit capability requirements.
After: inspect job history, usage, and signed receipt evidence.
Choose when it can accept work, which workloads it may run, and pause it at any time. Completed jobs appear in an itemized earnings ledger. New providers can connect a Mac computer now, but it does not receive paid work until the provider agreement review is complete.