Run your own AI models on your own servers. Lion AI Cloud organises Ollama and llama.cpp nodes behind one gateway with API keys, a queue and monitoring.
Screens captured from the actual software, filled with sample data.
Lion AI Cloud is the control plane LionHost uses for its own AI tools, such as the WHMCS support assistant. It connects servers running Ollama or llama.cpp and presents them as one gateway.
It is designed first for CPU servers, with no graphics cards required: the scheduler considers each node’s RAM, disk and load, and it is ready for GPU nodes when you need them.
Each application gets its own API keys, requests-per-minute limit, monthly quota, allowed models and pools, its own system prompt and a separate knowledge base. Requests enter a queue with priority classes, so customer support never waits behind a batch job.
CPU, RAM and disk telemetry from a read-only agent.
Pull, warm, delete and placement policies for models.
Limits, quotas and allowed models per application.
Background, batch, interactive, coding, support, critical.
Tokens/sec and latency history per node and model.
Maintenance windows, config backups, alerts and a firewall.
No. Lion AI Cloud is designed for CPU servers and also supports GPU nodes when available.
Ollama and llama.cpp.
On the servers you connect. Models run locally on your own nodes.