Skip to Alexandria
INFERENCE · INCLUDES ALEXANDRIA

PRODUCTION AI INFERENCE.ELITE MANAGED TOOLING.

Elastic or dedicated inference on one key. Alexandria ships with it — named tools, same wallet.

AthenaOlympusVesperHeraldAtlasSOON
import OpenAI from 'openai';
const client = new OpenAI({
baseURL: 'https://hypervize.tech/api',
apiKey: process.env.HYPERVIZE_API_KEY,
});
const response = await client.chat.completions.create({
model: 'google.gemini-3.7-flash' // or 'auto' — we pick the best model by default,
messages: [{ role: 'user', content: 'How deep is the Marianas Trench?'}],
max_tokens: 1024
});
// Enabled Alexandria tools inject on this same call
// Athena · Vesper · Herald · Ledger — same key, no second API
COMPUTE → REGISTRY → ENDPOINT

FROM TRAINING TO INFERENCE

  1. 01
    Fine-tune on Hypervize Compute — provisioned GPUs or HGX.Provisioned compute →
  2. 02
    Weights land in the Hypervize Registry. No USB-stick ritual.
  3. 03
    One-click dedicated endpoint. Same OpenAI-shaped API as elastic.Dedicated docs →
ZERO-CONFIG NETWORKING

BRING YOUR OWN DOMAIN.

Add a CNAME. We provision TLS and route to your dedicated endpoint. Auto *.hypervize.tech if you do not bring a name.

Dedicated docs →

CHOOSE YOUR DEPLOYMENT ARCHITECTURE

BEST FOR STANDARD MODELS

Elastic Inference (Burstable)

Priced per 1 Million Tokens

Curated frontier models on our shared cluster. No cold starts. You pay for the tokens you generate — not a junk catalog of a thousand IDs.

  • Auto-scaling concurrency
  • Fully managed
  • Global rate limits apply
BROWSE MODEL CATALOG
ENTERPRISE GRADE

Dedicated Endpoints

Priced per Hour (Compute Based)

Private GPU — fractional or full — with your fine-tune or any Hugging Face model. BYOD. Scale to zero. Throughput is yours, not a noisy neighbor's.

  • Custom weights
  • Guaranteed SLA & throughput
  • Unlimited tokens (hardware-bound)
DEPLOY CUSTOM ENDPOINT

Dedicated resources billed weekly based on usage (min 1h if used) + full-week addons via saved PM. Fine print on docs.

THE ALEXANDRIA PROJECT
KNOWLEDGE * MEMORY * ACTION

THE LIBRARY INSIDE THE PIPE

Enable a tool once. It is available on Chat, the API, MCP, and Chronos — same key, same prepaid ledger. Atlas is coming.

In Chat, the same library is the tools panel — not a second product.

Athena

Core
Foundational Intelligence

The essential brain add-on for every agent. Delivers real-time time/date awareness and safe basic calculations to ground reasoning and prevent hallucinations.

$0.00075 per call

Olympus

Core
Automatic Model Choice

Hypervize picks the right model for you — every time. On by default. No charts, no second-guessing Claude vs GPT vs Grok. Lock a model anytime.

Free

Desk

Core
Your Computer

Let Chat work with files on a computer you connect — Documents, Desktop, Downloads, Pictures. Pair in Chat. Saving, deleting, or running a command always asks first.

$0.015–$0.03 per call

Atlas

COMING SOON
Eternal Memory

The foundation of persistent agents. Store facts, histories, embeddings, and structured knowledge that survives every conversation, schedule, and session.

Storage-based • $0.002/MB + queries

Chronos

Automation
Autonomous Time

Enable scheduled tool and model recipes. When on: Schedule appears on chat replies that used tools, and API jobs work with your key. $0.01 per run plus tools and any model tokens.

$0.01 per run + tools + tokens

Vesper

Research
Web Intelligence

Full internet browsing for agents. Search or fetch any page and extract exactly what you need with natural language instructions. Supports structured output (JSON tables, metadata, headings, etc.).

$0.015 per call

Herald

Delivery
Direct Agent Voice

Send plain text, Markdown, or HTML emails to any address — only when you explicitly ask. Optional Library attachment via attachment_file_id (private docs loaded server-side). 20 emails/hour, max 10k char body.

$0.05 per call

Pandora

Actions
Open the Box

Full read and write access across your email, calendar, and file services. Open Pandora’s box and unleash your agents into your real digital world.

$0.015 per call

Iris

Creation
Image Creation

Turn language into images. Describe what you want; Hypervize selects Imagine Image, Imagine Image Quality, or Imagine Image 2.0 (precise). Pay per image — not a monthly creative suite.

$0.015 per call

Ledger

Files
Files & Data

Create, read, and convert files (CSV, JSON, PDF, text). Stored privately.

$0.03 per call
  1. 01 ENABLE
    In Alexandria (dashboard), or POST /api/tools with your key.
  2. 02 CALL
    /api/chat/completions or Chat. MCP execute is gated the same way; Chronos needs its own toggle.
  3. 03 RUN
    Tools run server-side. Billed as tool calls on the prepaid balance.

ELASTIC MODEL PRICING

Tokens and Alexandria tool calls draw from the same prepaid balance. Dedicated weekly invoices are GPU time only. Billing docs.

Model NameContext WindowInput (1M)Output (1M)
Gemini 3.1 Pro
google.gemini-3.1-pro
1M----
Kimi K3
moonshot.kimi-k3
1M----
Claude Fable 5.1
anthropic.claude-fable-5-1
1M----
Grok 4.6
xai.grok-4.6
500k----
GPT-5.6 Sol
openai.gpt-5.6-sol
1M----
GPT-5.5
openai.gpt-5.5
1M----
GPT-6 Astra
openai.gpt-6-astra
1.05M----