OOAICA

First 100 customers: $29 for a full yearfree 14-day trial

Stop renting your coding model.
Run it on hardware you own.

OAICA is a terminal AI CLI that points Claude Code, opencode and codex at a local Ollama or llama.cpp server, any OpenAI-compatible endpoint, or a machine you control. It hosts no models, needs no account, and never sees your code.

curl -fsSL https://oaica.com/install.sh | bash

Windows: irm https://oaica.com/install.ps1 | iex

14 days of the full paid set, no card. The free tier — local models, your own endpoints, the subagent tier — never expires.

Hosts no models of its own No account needed to run it 15 agent CLIs wired up 44 provider presets compiled in Three OSes, one installer Verifies offline — signed licence, no check-in

The part nobody prices in

What a coding agent costs you before it writes a line

Wiring one model into a terminal agent is a weekend. Wiring a fleet — tiers, context limits, failover, adapters — is the project. Here is the work OAICA already did, with a conservative estimate of what doing it yourself takes.

You would otherwise buildTypical effort
Endpoint config for every backend, with the same flags and env names across all of them4 h
Context-window probing and per-model token budgets, so long sessions stop truncating silently6 h
A health circuit breaker that moves a live stream to the other leg without dropping it10 h
Tier routing — one model to plan, another to execute — with the tool-call shapes normalised8 h
Per-request LoRA: loading, stacking, scaling, and keeping adapters isolated between callers12 h
Wrapping it all for Claude Code, opencode and codex, each with different expectations7 h
An installer for three operating systems that verifies its own download5 h
Then keep it working as the models keep changingongoing

Estimates for a competent engineer starting from nothing, at 2026 model APIs. Not a quote for your time — a list of what is already done.

See how it mixes providers → Skip to the install

Download

Pick your operating system. Each archive holds the oaica binary — unzip it, put it on your PATH, run it.

Three ways to run a model. One command in front of all of them.

Start with the machine under your desk. Move to a box on the network when you outgrow it. Nothing else changes.

Local models

Point it at a local Ollama or llama.cpp server on 127.0.0.1:11434. The common case needs no configuration at all.

oaica run llama3.2

Any endpoint

Set OAICA_HOST to any OpenAI-compatible base URL — a box on your network, a tunnel, or a hosted provider you already pay for.

export OAICA_HOST=https://<endpoint>

Claude Code ready

Run Claude Code, opencode and codex on your own model, with the context window and tiers configured automatically.

oaica launch claude --model <model>

Mix and match

Four tiers. Four different providers. One command.

Most launchers choose a vendor for you and stop there. OAICA runs a single session across every backend you have — a cloud model to plan, a local one to execute, a small one for the busywork, a long-context one for the moment the window fills — and names the leg that served each reply.

$ oaica launch claude \
    --model        <cloud-model>      # the vendor you already pay
    --sonnet-model <local-model>      # the box under your desk
    --haiku-model  <small-local>      # titles, topic detection, busywork
    --oversize     <long-context>     # takes over when the window fills
    --route-policy weighted \
    --shard <local-model>:3 --shard <cloud-model>:1

The model names are yours — OAICA ships no fixed list and probes whatever your endpoint actually serves. Walk it once with the wizard, save it as a named plan, and every later launch is one word:

$ oaica launch claude --plan work

Four tiers, any mix

A primary, an execution tier, a background tier and an oversize/compaction tier. Nothing forces them onto the same provider — or onto the same machine. Each tier resolves from a flag, a saved plan, or a standing config, in that order.

Six route policies

local-first (the default), local-only, remote-first, remote-only, auto and weighted. weighted splits healthy traffic by weight on a session-sticky hash, so a conversation keeps its prefix cache instead of bouncing between legs.

Route one tool, one way

The turn that digests a tool's output is the expensive one. oaica config set tool-model-default secondary moves every tool-heavy turn to another leg; oaica config set tool-model Bash tertiary does it for a single tool.

See what the paid tiers add →

Built for real sessions

The features that matter once a model is running your work

Not a chat box. The things that go wrong on day three, handled.

Plan with one model, execute with another. Point Opus-tier planning and Sonnet-tier execution at different backends — a cloud model and a local one, together.
Failover that never breaks a stream. A health circuit breaker moves a conversation to its other leg; X-Oaica-Route names which one served each reply.
Per-request LoRA. Toggle, stack and scale adapters for your session only — isolated from every other caller of the same model.
Build a website from a sentence. oaica site new plans, writes, sanitizes and assembles a static page you can host anywhere.
Read-only health checks. oaica doctor probes every configured leg and exits non-zero on any failure — safe to run anywhere.
No account, no hosted models. OAICA runs on your machine and talks only to the endpoints you configure. Your code never leaves it.

Check for yourself

Things a terminal launcher does not usually do

Every claim below is a command you can run on the free tier. No asterisk, no "coming soon".

Four model tiers in one launch. --model, --sonnet-model, --haiku-model, --oversize — four providers if you have four, and the wizard saves the whole arrangement as a named plan.
Six route policies, chosen per launch. --route-policy local-first|remote-first|auto|local-only|remote-only|weighted.
Weighting that respects a session, not a round-robin. --shard <model>:3 --shard <model>:1 — a consistent hash keeps one conversation on one leg, so its prefix cache stays warm.
A model of its own for one tool. oaica config set tool-model Bash secondary sends the turn that digests a huge tool output to the cheap leg, and leaves the rest alone.
Anthropic-shaped integrations meet OpenAI-shaped providers at loopback. The launcher runs the translation locally and hands the integration a 127.0.0.1 address, so your provider key never leaves the machine and the tool never sees the raw stream.
Probing instead of a fixed catalogue. The context window is measured on the live endpoint before a session starts, so a number written in a catalogue file never overrides what the model can actually take. 44 provider presets compile into the binary; oaica remote add mine --base-url … --api-key-env … adds yours, and there is no release to wait for.
A licence that verifies with no network. Your rights travel as an Ed25519-signed claim, verified against a key pinned in the binary — so a paid licence keeps launching on a plane, and not one byte of your code goes anywhere to prove it.
One binary, 15 agent CLIs. claude, codex, cline, copilot, opencode, qwen, droid, pi and the rest of them, all pointed at the model you chose.
A support bundle that redacts itself. oaica doctor probes every configured leg and exits non-zero on failure; --report writes the same thing with credentials stripped.
No metering of what you run through it. Your inference bill, if you have one at all, goes to whoever serves the model. OAICA is not in the path and does not bill by the token.

Download it and check — free

Doing it yourself, or running OAICA

Both are legitimate. One of them is a weekend; the other is every weekend.

A hand-rolled wrapper

  • One backend, and a second one that half-works
  • Context limits found by watching output truncate
  • A dropped stream is a lost conversation
  • Every model swap is a small refactor
  • LoRA adapters leak between sessions
  • You are the maintainer, permanently

OAICA

  • Ollama, llama.cpp, vLLM, any OpenAI-compatible URL
  • Context window probed and budgeted per model
  • Circuit breaker moves the stream, mid-reply
  • New model: change one flag, keep your setup
  • Adapters scoped to your session, nobody else's
  • Someone else ships the fixes

Skip the weekend — start free for 14 days

Pricing

Pay for the year, not the tokens

One licence, three devices, and no metering of what you run through it. Your inference bill — if you have one at all — goes to whoever serves the model, not to us.

Free tier

$0

Never expires. No card, no account.

  • Local Ollama and llama.cpp models
  • Your own OpenAI-compatible endpoints
  • The subagent tier
  • oaica doctor and the installer
Download free

Code Monthly

$7 / month

Cancel any time.

  • Everything in Code Annual
  • No yearly commitment
  • Same 14-day trial first
  • Three devices per licence
Get Code Monthly

Every fresh install runs the full paid set free for 14 days, no payment details. Full pricing and what is in each tier →

Earn 30% of every payment, for as long as they stay

If you write about running models on your own hardware, you are the audience. Join from your dashboard, get a link, and keep 30% of each subscription payment it brings in — recurring, not a one-off bounty.

See the program

Questions worth asking before you pay

Does my code go to you?

No. OAICA hosts no models and has no inference servers. It is a client that talks to the endpoints you configure — usually 127.0.0.1. There is nothing in the paid tier that sends your code anywhere else.

What does the licence actually unlock?

The paid tiers add what a single local model cannot do on its own: routing one tier of work to one model and another tier to a second, per-request LoRA adapters, failover that survives a dropped stream, the Claude Code / opencode / codex launchers, and the site builder. The free tier keeps local models, your own endpoints and the subagent tier, for ever.

What happens after the 14 days?

Nothing breaks. The paid features stop responding and the free tier carries on working, unchanged and unlimited in time. You can buy a licence at any point afterwards and the paid features come back on the same install.

Is the $29 a subscription?

It is the first year of a yearly subscription. The launch price covers the first year only; from the second year it renews at $59 unless you cancel. Code Monthly is $7 a month if you would rather not commit to a year.

How many machines can I run it on?

Three devices per licence. Deactivate one from your dashboard to move it to another — there is no waiting period.

Which models does it work with?

Anything that speaks the OpenAI API: Ollama, llama.cpp, vLLM and most hosted providers. The launchers are tested against the models people actually run locally, and oaica doctor tells you what a given endpoint supports before you start a session on it.

What if I want a refund?

Ask within 30 days and you get it, no interrogation. Email oaica@sprapp.com, or use the support button — the refund path is the same one a human reads.

Where do I get help?

The docs cover install, configuration and every command. There is a support chat on the site and at oaica@sprapp.com — both reach the same inbox.

Get Code Annual — $29 first year

Be honest about whether this is your problem

A tool you do not need is a subscription you have to remember to cancel. Here is the shortest version of who should close this tab.

Not for you if
  • You are happy renting a hosted coding model and never want to think about hardware. Pay the vendor directly; you are already getting what you want.
  • You want a chat window. This is a command line, and it wires other people's agents to your models — it is not a nicer ChatGPT.
  • Your machine is managed so you cannot run a local server or set an environment variable. The free tier still works, but the reason to buy does not.
  • You want a tool that locks itself the second a payment fails. OAICA honours the period you paid for and verifies offline, which is the opposite policy.
Built for you if
  • You already have one good model — local, rented, or on a box at the office — and want a coding agent pointed at it.
  • You have more than one backend and are tired of hand-editing endpoint config to switch between them.
  • You have been bitten by a stream dying mid-answer, or by a context window you only discovered when output started truncating.
  • You want to pay for a year of software, not for tokens you cannot audit.

What you are actually buying

Everything in the paid set. $29 for the first year.

Not a lighter tier — the whole product, on three devices, with nothing metered.

  • Tier routing and saved plans — four tiers, any mix of providers, one word to relaunch
  • Stream-preserving failover — a health circuit breaker that moves a live reply to the other leg
  • Per-request LoRA — stacked, scaled, and scoped to your session alone
  • Fifteen agent CLIs — Claude Code, codex, opencode and twelve more, one command in front of all of them
  • The site builder — plan, write, sanitize and assemble a static page from a sentence
  • Three devices, with updates and support while your subscription is paid

The risk is not yours. Run the entire paid set free for 14 days first, with no card. If it turns out not to be for you, ask within 30 days and you get your money back.

The price does move. The first 100 customers keep a full year at $29; after that it renews at $59. $7 a month if you would rather not commit to the year.

Still not sure? Compare the editions line by line, or run oaica doctor after install and let it tell you what your setup can do.