No bill for every request
Your model runs on a computer you already own, so you aren't paying an inference provider for each request. No OpenAI, Anthropic, or other paid inference service is required.
Stratumus connects local AI models on computers you own to one simple, private workspace. Install the Connector, pair your Mac, and use oMLX or Ollama without paying an inference provider for every request.
This is the installation experience currently being productized for early access.
Open one macOS app. No Node.js, ports, or configuration files.
Approve the connection from stratumus.com.
Stratumus discovers oMLX, Ollama, and compatible installed models automatically.
A prompt entered at stratumus.com travels through an authenticated relay to your selected Mac. Your local model generates the response, and Stratumus streams it back to your browser.
Want the details of what each step can see? How your data moves.
The Stratumus web app or PWA, wherever you are.
Your deviceSigns you in, keeps your conversations, and relays requests to your Node.
Hosted by StratumusRuns on your Mac and connects out to the control plane.
Your MacRuns the model you picked, on your hardware.
Your MacStratumus Nodes use a common provider interface, allowing the same chat experience to work with oMLX, Ollama, and future local engines.
Stratumus discovers supported providers and installed models on each connected Mac. Choose a primary Node, arrange fallbacks, and see the health of your private AI system at a glance.
Your selected model runs on your Stratumus Node, using hardware you own. Stratumus securely relays requests between the web app and your Node so you can use it from anywhere. No OpenAI, Anthropic, or other paid inference service is required.
Inference is local. Requests travel through a relay. We spell out both.
Prompts and responses pass through the Stratumus control plane, currently hosted on Railway, on their way between your browser and your Node. Your browser connects over an encrypted connection, and your Connector connects outbound through an authenticated relay.
Conversations are managed by the Stratumus application, which is how your history persists between sessions.
Your models and inference. The selected model runs on your Node, and no third-party inference service is involved.
Web search is optional. When you turn it on, your search query goes to an external search service to fetch results.
Direct or end-to-end encrypted transport. Until then, requests are relayed through the control plane as described above.
Your model runs on a computer you already own, so you aren't paying an inference provider for each request. No OpenAI, Anthropic, or other paid inference service is required.
Run the models and runtime you prefer, and keep the same chat experience when you change them.
Your Node keeps working when you're not home. Use it from a browser or the Stratumus PWA, wherever you are.
Stratumus is starting on macOS with oMLX and Ollama. Leave your email and we'll let you know when early access opens.