Your own inference control plane
A private AI inference control plane that runs local large-language-model drives on your own hardware.
SpazHost provides a secure, private environment for running large language models locally. Instead of paying per-token cloud APIs, you run your own models on your own hardware with full control over access, routing, and performance.
Model Drives
Spin up and hot-swap local models with lane-based routing. Run multiple models simultaneously and switch between them seamlessly.
Typed Models
Pick a model by what it does — chat, vision, document OCR, embeddings, image generation. Request routes straight to the right drive.
OpenAI-Compatible Gateway
Drop-in replacement for OpenAI's /v1 API. Integrate with existing applications without code changes.
Access Control
Per-key, per-lane access control. Secure your models with granular permissions and usage limits.
Local Admin
Drive, GPU, and container management runs in a local app over SSH — never a public login. Your control plane stays on your machine.
Own Your Stack
Runs entirely on your hardware. No per-token cloud bills. Keep your data and compute private.
Available models
Grouped by type. Live availability updates from the host.
Chat
General reasoning and coding models.
Vision
Image understanding and multimodal Q&A.
Document OCR
Extract text, tables, and structure from documents.
Embeddings
Vector embeddings for search and RAG.
Image Generation
Text-to-image on local hardware.
Model list is served from the host's catalog feed (public-safe: type, context, online status — no drive/hardware detail).
Ready to run your own inference?
Get a key