Ninja for enterprise

The next billion employees will be AI employees

An unmetered, unlimited-intelligence AI workforce that runs 24/7 inside your own cloud. It finishes work end to end, builds endless automation, self-evolves to your business, and collaborates with your team in real time.

SaaS · VPC · On-prem · Air-gapped · Palo Alto, CA

Why Ninja Enterprise

Your cloud. Your models. Unmetered.

Fully deployed in your VPC, on any cloud

Ninja runs entirely inside your own virtual private cloud on Azure, AWS, or GCP, or on-prem, or fully air-gapped. Your data, tools, and models stay behind your firewall. Run open-source models on your own GPUs and nothing leaves your perimeter; choose a frontier model and the only outbound call is to the endpoint you pick, under your terms — never pooled, never trained on.

Unmetered intelligence on your own GPUs

Run open-source models on dedicated GPUs and make Ninja unlimited. No per-seat fees, no per-token meter. You pay the hourly GPU cost plus our software license per GPU node: a flat line item that counts toward your existing cloud commitments. 5-10x lower cost than metered frontier seats at scale.

The economics

A flat line item, not a runaway meter.

Metered per-seat and per-token pricing punishes you for using AI more. Dedicated-GPU deployment is fixed-cost: the more your AI workforce does, the more you save per unit of work.

Predictable budget.

One flat cost per GPU node, forecastable to the dollar.

Funded by cloud commits.

Counts toward Azure MACC, AWS EDP, and GCP CUD — no net-new budget.

Scales without penalty.

Add work, not invoices. Usage is capped by hardware, not by a meter.

Architecture

One self-contained stack, in your environment.

Ninja deploys as one self-contained stack inside your cloud or on-prem — not a constellation of outbound SaaS calls. Everything an agent needs runs in the box you control.

LiteLLM is the control plane

Every model, MCP tool, guardrail, and spend log is registered and governed in one place — so security reviews one surface, not fifty.

Caddy is the single entry point

An integration gateway maps MCP to your apps; each agent runs in its own isolated sandbox with a real browser (Phantoms).

Identity via OIDC single
sign-on

Tokens encrypted at rest (Fernet); an in-stack git server (Gitea) so code and learnings never leave the box. Backing services — Postgres, Redis, the message queue, Prometheus + Grafana for audit — all stay inside.

The only outbound call is to the model you choose

Run open-weight models (Kimi K2/K3, GLM 5.2) on your own NVIDIA GPUs and there is no egress at all; point at a frontier model and that endpoint is the single, auditable exit.

Security

Governed at the gateway. Isolated by workspace

Ninja is designed for regulated, high-trust environments, deployed where your security team can see and control it.

Inside your perimeter

VPC, on-prem, or fully air-gapped. A dedicated, isolated VM per agent and thread.

Your controls

SSO/SCIM, RBAC, and full audit trails. Encrypted in transit and at rest. Because Ninja runs inside your own perimeter, your existing SOC 2 and HIPAA controls apply to it. SOC 2 Type II in progress; HIPAA-eligible deployments inside your BAA-covered environment.

Your data, forever

Your data never trains anyone’s model. PHI and sensitive data stay inside your compliance boundary.

Human-verified

A person confirms high-impact actions, with a full audit trail and a live view of the agent’s browser. We show the controls; we don’t claim zero hallucination. For regulated use, agent output still requires human verification.
HIPAA-eligible
SSO/SCIM
RBAC
Audit trails
SOC 2 Type II, in progress

Deployment path

Prove it in a pilot. Own it in your cloud.

Start managed, prove value on a real workflow, then move everything behind your firewall.

Phase 1

Prove it
Manage SaaS Pilot

Your AI workforce live in your Slack within a week. Credit-based, fully managed, day-one ready. Prove value on a real workflow first.
Phase 2

Manage it
Deployed in your VPC

Move to a fully-owned deployment on dedicated GPUs. Unmetered open-source models, no credits, no per-seat fees. Software license per GPU node: one flat line item.

Models

The best models. Any cloud. AI you own.

Run the best frontier and open-source models, swap any time, and upgrade as new ones ship. No one else lets you choose across all of these, inside your own cloud.

Frontier

via the gateway
Claude Opus 5
Claude Fable 5
GPT-5.6

Open-weight

self-hosted on your GPUs
Kimi K3
GLM 5.2
Swap anytime

Bring your strictest reviewer

We deploy behind your firewall, prove it on your hardest workflow, then you scale. Let’s find the first one.