Elastic & Dedicated Inference on HyperVize Compute

PRODUCTION AI.ZERO DEVOPS.

Access industry-leading models via our elastic inference API, or deploy your own custom fine-tunes to dedicated, auto-scaling endpoints in one click. Part of the full HyperVize platform including bare metal clusters and provisioned compute for training. Integrate instantly using our high-performance native SDKs or REST API.

VIEW MODEL PRICING
OpenAI Drop-in Replacement
import OpenAI from 'openai';
const client = new OpenAI({
baseURL: 'https://hypervize.tech/api',
apiKey: process.env.HYPERVIZE_API_KEY,
});
const response = await client.chat.completions.create({
// Use our elastic models OR your dedicated endpoint ID
model: 'moonshot.kimi-k3',
messages: [{ role: 'user', content: 'How deep is the Marianas Trench?'}],
max_tokens: 1024
});
Deployment Pipeline
Active
01
HypervizeTM Compute
Fine-tune Llama 4 on 8x B200 HGX
Save weights to HypervizeTM Registry
02
Dedicated Inference
1-Click deploy to highly available endpoint

FROM TRAINING TO INFERENCE IN SECONDS.

Don't waste days configuring vLLM, load balancers, and custom domains. Once your fine-tuning job finishes on our Compute tier, you can instantly mount your custom weights to a Dedicated Inference Endpoint.

  • Custom Routing: Auto-assigned *.hypervize.tech subdomain out of the box, with full BYOD support available.
  • Private Weights: Your proprietary data never leaves our secure fabric.
  • Scale to Zero: Configure idle timeouts to minimize infrastructure costs.
Zero-Config Networking

BRING YOUR OWN DOMAIN.

Stop wrestling with NGINX, reverse proxies, and Let's Encrypt. HyperVize automatically provisions secure, load-balanced API endpoints for your models the moment they launch. Use our auto-generated subdomains, or map your own custom domain in seconds.

Seamless BYOD Routing

Simply add a CNAME record. We handle the ingress routing directly to your dedicated vLLM instance.

Automated TLS / SSL

Enterprise-grade encryption out of the box. Certificates are automatically provisioned and renewed seamlessly.

ENDPOINT SETTINGS
ROUTING ACTIVE
Mounted Weights
llama3-70b-finetune-v2
Private
Custom Domain Mapping (BYOD)
CNAME record verified
SSL Certificate Generated
Base URL Ready
https://api.neural-dynamics.com/v1

CHOOSE YOUR DEPLOYMENT ARCHITECTURE

Whether you need the absolute lowest cost per token for standard models, or guaranteed throughput for your custom fine-tunes, we have a tier for you.

BEST FOR STANDARD MODELS

Elastic Inference (Burstable)

Priced per 1 Million Tokens

Access our curated selection of the world's best models (Kimi K3, Claude Fable 5, GPT-5.6 Sol, Grok 4.5) hosted on our highly-available global cluster. We don't give you thousands of options. We give you the ones that work well. Zero cold starts. You pay strictly for the tokens you generate.

  • Auto-scaling concurrency
  • Fully managed by HypervizeTM
  • Global rate limits apply
BROWSE MODEL CATALOG
ENTERPRISE GRADE

Dedicated Endpoints

Priced per Hour (Compute Based)

Provision a dedicated fractional or full GPU instance loaded with vLLM. Perfect for hosting your own custom weights with Bring-Your-Own-Domain (BYOD) support and zero noisy neighbors.

  • Host custom fine-tuned weights
  • Guaranteed SLA & Throughput
  • Unlimited Tokens (Hardware bound)
DEPLOY CUSTOM ENDPOINT

Dedicated resources billed weekly based on usage (min 1h if used) + full-week addons via saved PM. Fine print on docs.

ELASTIC MODEL PRICING

Model NameContext WindowInput (1M)Output (1M)
Kimi K3
moonshot.kimi-k3
1M----
Claude Fable 5
anthropic.claude-fable-5
1M----
Grok 4.5
xai.grok-4.5
500k----
GPT-5.6 Sol
openai.gpt-5.6-sol
1M----
GPT-5.5
openai.gpt-5.5
1M----
GPT-5.5 Pro
openai.gpt-5.5-pro
1M----