The Cost-First AI Architecture: Run On-Premise by Default, Pay for the Cloud Only at the Peaks
    Cost Analysis
    Hybrid
    On-Premise
    Architecture

    The Cost-First AI Architecture: Run On-Premise by Default, Pay for the Cloud Only at the Peaks

    10 juin 2026Oliver Glas

    The Cost-First AI Architecture: Run On-Premise by Default, Pay for the Cloud Only at the Peaks

    Most teams treat AI infrastructure as an either/or decision. Run it in the cloud and accept a bill that scales with every request, or run it on your own hardware and accept a hard ceiling on capacity. Both framings leave money on the table. There is a third option that captures the economics of owned hardware for everyday load while keeping the elasticity of the cloud for the moments you actually need it — and KeepUrAi is built to run exactly this way.

    This guide shows how to deploy the platform so that the cheap resource does the heavy lifting and the expensive resource is billed only when demand spikes.

    The Economic Insight

    The cost of running AI is dominated by one thing: GPU inference. Everything else — the web tier, the API layer, the database — is inexpensive and rarely the bottleneck. So the entire cost question reduces to where the GPU work happens.

    Two facts make the decision straightforward:

    • On-premise GPU has a near-zero marginal cost. Once the hardware is paid for, each additional request is essentially free. Electricity is the only variable.
    • Serverless cloud inference has a near-zero idle cost. Providers like Ollama's cloud and Comfy's cloud API bill per token or per job. When you send nothing, you pay nothing.

    Put those together and the optimal strategy writes itself: let owned hardware absorb the baseline load at almost no marginal cost, and spill only the overflow to pay-per-use cloud inference during peaks. You never pay for an idle cloud GPU, and you never queue users for hours when your own GPU is full.

    The Architecture at a Glance

                        ┌─────────────────────────────────────────┐
                        │            ON-PREMISE (baseline)         │
       Users  ─────────▶│  Web + API + Gateway                     │
                        │  PostgreSQL + Vector Store (all data)    │
                        │  Local GPU: Ollama + ComfyUI             │
                        └───────┬───────────────────────┬─────────┘
                                │                        │
              overflow only    │                        │  premium work only
              at GPU saturation▼                        ▼  (deliberate escalation)
            ┌───────────────────────────┐   ┌───────────────────────────────┐
            │  CLOUD (elastic, per-use) │   │  CLAUDE CODE (frontier agent) │
            │  Ollama cloud inference   │   │  Connects via MCP, no GPU      │
            │  ComfyUI cloud generation │   │  Draws on on-prem knowledge    │
            └───────────────────────────┘   └───────────────────────────────┘
    

    The application stays entirely on-premise. Your documents, your conversations, your user data — none of it relocates. The only things that ever cross outward are an individual inference call (when your local GPU is busy) and the specific context a premium task needs (when you deliberately escalate it).

    Step 1: Deploy the Full Stack On-Premise

    KeepUrAi ships as a complete, self-contained stack and supports three deployment paths — choose whichever matches your environment:

    • Debian / RPM packages with systemd, for a traditional on-prem server install
    • Docker Compose, for a single-host containerized deployment
    • Kubernetes (Helm), for a clustered deployment you manage yourself

    All three install the same components: the backend API, the web frontend, the mobile gateway, PostgreSQL with the vector store, and the local AI services. This is your baseline tier, and for most organizations it will handle the large majority of day-to-day traffic on hardware you already control.

    Step 2: Keep Your Data Where It Is

    The data layer is a modest, GPU-free machine: PostgreSQL with the pgvector extension, holding both your relational data and your retrieval-augmented-generation (RAG) embeddings. It runs comfortably on commodity hardware with an NVMe disk and a sensible amount of RAM.

    Crucially, this tier never needs to move or replicate for the architecture to work. Because only inference calls overflow to the cloud — not data — you keep a single database on-premise. That eliminates the most expensive and error-prone part of most hybrid designs: cross-site database synchronization.

    Step 3: Configure the Cloud as Your Overflow Target

    The platform already understands two operating profiles — local (your own GPU) and cloud (external inference endpoints) — and exposes both through environment variables. You point the cloud profile at the serverless endpoints you want to use for overflow:

    # Local baseline (default) — your own GPU does the work
    SPRING_AI_OLLAMA_BASE_URL=http://localhost:11434
    COMFYUI_MODE=local
    
    # Cloud overflow target — billed only per request
    SPRING_AI_OLLAMA_BASE_URL=<ollama-cloud-endpoint>
    SPRING_AI_OLLAMA_CHAT_OPTIONS_MODEL=<cloud-catalog-model>
    COMFYUI_MODE=cloud
    COMFYUI_CLOUD_API_KEY=<your-cloud-api-key>
    

    Because the platform routes all AI traffic through its own backend — frontends and mobile clients never call an AI service directly — you have a single, central place to decide whether a given request is served locally or in the cloud. That central control point is what makes a clean overflow policy possible.

    Step 4: Let the Queue Absorb Load Before You Spend

    KeepUrAi's chat pipeline is asynchronous by design. Instead of holding a connection open while a model thinks, the client submits a request and polls for the result:

    POST /ask/submit      → returns a job id immediately
    GET  /ask/status/{id} → returns progress and the answer as it streams
    

    This matters for cost. Because users are already polling, your local GPU queue can absorb a meaningful amount of load before anyone experiences a delay worth avoiding. The practical rule is simple: let on-premise queue the work first, and overflow to the cloud only when the wait would become unacceptable. Queueing on owned hardware is free; overflowing costs money. The async path lets you exploit that difference instead of paying to avoid every few seconds of latency.

    Step 5: Set Cost Guardrails

    A cost-first architecture deserves cost-first controls. Three policies keep the bill predictable:

    | Control | Purpose | |---|---| | Conservative overflow threshold | Spill to the cloud only when the local GPU is genuinely saturated, so cheap hardware does the maximum possible work. | | Daily cloud-spend cap | Past a budget ceiling, stop overflowing and queue on-premise instead — trading a little latency for a hard cost limit. | | Conversation-level affinity | Once a conversation overflows to the cloud, keep its remaining turns there, rather than switching back and forth mid-thread. |

    Together these ensure the cloud is a relief valve, not a default — and that a traffic spike can never quietly turn into a runaway invoice.

    Add a Premium Capability Tier with Claude Code

    The two tiers so far differ only in where inference runs. The same principle — pay only for what the request demands — extends to how capable the engine needs to be. Not every task is equal: most are routine, a few are genuinely hard. KeepUrAi lets you connect Claude Code, Anthropic's agentic coding tool, as a deliberate, high-capability tier for the work that justifies it.

    Claude Code connects to the platform's built-in agent interface (MCP), authenticated like any other connected agent and routed through the same gateway as the rest of your traffic. It requires no GPU and no provisioning of its own — it is an agent your developers already run, joined to your private hub. Once connected, it can:

    • Search your on-premise documents and knowledge base through scoped, audited tool calls
    • Generate images and translate text using the platform's services
    • Pick up tasks that users submit from the web or mobile app, complete them, and return the results

    This rounds out a three-tier cost ladder, each rung billed in proportion to the value it delivers:

      Tier 1  Local GPU        ──▶  baseline chat, RAG, search        (near-zero marginal cost)
      Tier 2  Cloud overflow   ──▶  same work at peak load            (pay-per-token, spikes only)
      Tier 3  Claude Code      ──▶  complex, agentic, high-value work (per-use, invoked deliberately)
    

    The cost logic is the same as the overflow valve: routine load never reaches Tier 3, so it stays inexpensive; only the hard, valuable tasks escalate to it. And because Claude Code draws on your private knowledge through controlled tool calls rather than a wholesale data export, you combine frontier-grade reasoning with your own on-premise context — without buying frontier-grade hardware to host it.

    Keeping the Experience Consistent

    There is one trade-off to plan for honestly. Work served by the cloud overflow tier — or by the Claude Code tier — is handled by a model that is not byte-for-byte identical to your local one. For most general assistance this difference is minor, but it is worth managing:

    • Choose a cloud catalog model whose behavior is close to your local one.
    • Use conversation-level affinity (Step 5) so a single conversation never changes voice midway through.
    • Reserve overflow for interactive, latency-sensitive requests, and reserve the premium tier for tasks whose difficulty justifies it; let batch and background work wait in the on-premise queue where consistency and cost both favor staying local.

    When Privacy Is the Priority Instead

    This guide optimizes for cost on the assumption that sending occasional inference — or a deliberately escalated task — to an external provider is acceptable. When data residency or confidentiality is the overriding concern, the same platform runs in a fully on-premise profile with no external calls at all: every token stays on your hardware. The architecture is unchanged; you simply leave both outward valves closed — the cloud overflow tier and the external agent tier alike — and run entirely on Tier 1. The point is that you decide where the line sits, per deployment, rather than the platform deciding for you.

    A Starting Checklist

    Before going live with cost-first overflow, confirm:

    • [ ] The full stack is deployed and serving traffic from on-premise hardware
    • [ ] PostgreSQL with the vector store holds all application data locally
    • [ ] Local Ollama and ComfyUI handle the baseline load
    • [ ] Cloud inference endpoints are configured and reachable as the overflow target
    • [ ] An overflow threshold is set conservatively against local GPU saturation
    • [ ] A daily cloud-spend cap is in place
    • [ ] Conversation-level affinity is enabled for overflowed sessions
    • [ ] Claude Code is connected as an authenticated agent, scoped to the work that justifies it
    • [ ] Monitoring shows the local / cloud / premium request split and spend over time

    Conclusion

    The cheapest elastic AI architecture is not a second data center in the cloud, and it is not a load balancer that flips your whole application to a more expensive home. It is a single, mostly on-premise deployment that does its everyday work on hardware you already own, reaches for pay-per-use cloud inference only at the peaks, and escalates to a frontier agent only for the work that earns it — paying for spikes and for hard problems, never for idle capacity.

    KeepUrAi is designed for precisely this shape: one self-hosted stack, your data kept local, local and cloud inference profiles built in, Claude Code connectable as a premium capability tier, and the central control needed to choose between them request by request. If you are weighing what a cost-first hybrid deployment would look like for your own workload, we are happy to walk through it with you.