Your own inference control plane

A private AI inference control plane that runs local large-language-model drives on your own hardware.

SpazHost provides a secure, private environment for running large language models locally. Instead of paying per-token cloud APIs, you run your own models on your own hardware with full control over access, routing, and performance.

Model Drives

Spin up and hot-swap local models with lane-based routing. Run multiple models simultaneously and switch between them seamlessly.

Typed Models

Pick a model by what it does — chat, vision, document OCR, embeddings, image generation. Request routes straight to the right drive.

OpenAI-Compatible Gateway

Drop-in replacement for OpenAI's /v1 API. Integrate with existing applications without code changes.

Access Control

Per-key, per-lane access control. Secure your models with granular permissions and usage limits.

Local Admin

Drive, GPU, and container management runs in a local app over SSH — never a public login. Your control plane stays on your machine.

Own Your Stack

Runs entirely on your hardware. No per-token cloud bills. Keep your data and compute private.

Available models

Grouped by type. Live availability updates from the host.

Chat

General reasoning and coding models.

Vision

Image understanding and multimodal Q&A.

Document OCR

Extract text, tables, and structure from documents.

Embeddings

Vector embeddings for search and RAG.

Image Generation

Text-to-image on local hardware.

Model list is served from the host's catalog feed (public-safe: type, context, online status — no drive/hardware detail).

Ready to run your own inference?

Get a key