ucloud global logo

Build Your Own AI Model Platform

White Label AI Gateway — Build your own OpenRouter.

One API for 200+ AI models, with white-label branding and built-in user, access, and billing management.

AI Model Aggregation
Full White Label
AI Service Monetization
Security & Compliance
Your Users & Apps
Employees
Internal users
Business Systems
Enterprise apps
AI Agents
Autonomous agents
Online
UToken
AI Gateway
Smart RoutingSafety GuardrailsMetering & Billing
Models & Providers
Self-Hosted Models
DeepSeekGLMQwenKimi
Model Services
UCloud AstraFlow
Leading Model APIs
GPTClaudeGemini

Why UToken

AI Token governance for private enterprise use cases. From access to governance, UToken makes model calls unified and manageable.

01

One gateway for multiple models, less integration work.

Unified access to local, cloud, and third-party models. UToken is a unified AI Token gateway with access control, metering, and security.

  • Unified Access: One API for local, cloud, and third-party models.
  • Less Adaptation: No separate integrations for systems, agents, or apps.
  • Fast Reuse: Quick access for enterprise apps and assistants.
Web AppAI AssistantAgentBI System
UToken Gateway
  • Smart routing
  • Safety checks
  • Metering
Local LLMCloud LLMThird-party LLM

Unified access replaces point-to-point integrations.

02

Pre-call Risk and Budget Control

Centrally manage permissions, quotas, and frequency to stop risks before calls are made. UToken supports API Keys, IP whitelists, model access control, call limits, and budget caps for secure operations and cost control.

  • Access control: by tenant, project, customer, and API Key.
  • Rate limiting: QPS, TPS, and TPM limits to prevent traffic spikes.
  • Budget control: daily and monthly caps to avoid overspending.
Access Policy Panel
RoleAdmin · Tenant AActive
API Keyak-ut-****9f2cActive
IP Whitelist10.24.0.0/16Active
Model AccessDeepSeek-*, Qwen-*Active
Rate Limit60 QPS · 400 TPMActive
Budget Cap$500 / monthActive

Permissions, amounts, and frequency, all under control.

03

Intelligent Routing & Auto-Switching: Balancing Performance and Cost

Auto-match requests to the right models, avoiding wasting high-performance resources on simple tasks. UToken intelligently routes requests to the most suitable models, optimizing overall efficiency.

  • Auto-Selection: Basic models for simple Q&A; high-performance models for complex reasoning and Agents.
  • Better Performance: KVCache-aware scheduling routes to cache-hit replicas, reducing first-token latency.
  • Higher Throughput: Handle more concurrent requests on the same hardware, maximizing compute efficiency.
Incoming requestsUToken Smart Routing
Basic modelsSimple Q&A
High-performance modelsComplex reasoning
Private modelsCompliance tasks
Third-party modelsOverflow & fallback

Automatic routing, better cost-performance.

04

Metering & Cost Splitting: Transparent Expenditure

Go beyond Token counts; accurately track users, projects, and specific costs. UToken provides Token distribution, quotas, stats, billing, and settlement.

  • Usage Tracking: Monitor consumption by tenant, customer, project, and API Key.
  • Cost Accounting: Support metering to assess the real expenses of models and businesses.
  • Clear Splitting: Enable split settlements for multi-level scenarios like enterprises, channels, and ISVs.
Tokens by Team
R&D$1,346.5
Customer Support$833.1
Marketing$576.8
Finance$448.0
Total this month$3,204.4

Per-request metering, clear cost attribution.

05

Full-Link Security Audit & Local Data Storage

End-to-end traceability from requests to session logs. Fully traceable and auditable process from requests to sessions to meet compliance needs.

  • Unified Auth: Centralized access via UToken gateway to eliminate direct model connection risks.
  • AI Guardrails: Enable blocking, desensitization, and logging to prevent prompt attacks and data leaks.
  • Local Storage: Store prompts, responses, and context locally to ensure data sovereignty.
Request
Auth
Guardrails
Model Call
Response Log
Local Storage
Today 10:24:07 · tenant-a · ak-ut-9f2c · deepseek-v3 · 1,204 tokens · allowed
Today 10:24:09 · tenant-b · ak-ut-31d0 · qwen-max · 856 tokens · desensitized
Today 10:24:12 · tenant-a · ak-ut-9f2c · deepseek-v3 · 2,011 tokens · allowed

Every call traceable, data stays local.

UToken: Enterprise-Grade AI Token Governance Platform

Unified AI Token Gateway: connecting local inference, cloud compute & multi-vendor models. UToken manages underlying heterogeneous computing power, delivering secure, highly available, and operable AI service governance at the upper layer.

Application Layer

  • Enterprise Employees
    Corporate staff & office users
  • Internal Systems
    Business apps & platforms
  • AI / Coding Agents
    Autonomous agents & copilots
  • Channel Partners
    ISVs & ecosystem partners

UToken AI Gateway

Unified Access · Governance · Metering
  • Unified AuthCentralized authentication
  • Smart RoutingAuto model switching
  • AI GuardrailsBlocking & masking
  • Metering & BillingUsage & cost accounting
  • Quota & BudgetRate & budget control
  • Multi-tenant ManagementTenant isolation

Heterogeneous Compute & Models

  • Local Inference
    • Enterprise GPU Clusters
    • UCloudStack
    • Local LLMs
  • Cloud Compute
    • AstraFlow
    • Cloud GPU
    • 3rd-party Compute
  • Multi-Vendor Models
    • MaaS Platforms
    • Industry Models
    • OpenAI / Anthropic

One governance layer over heterogeneous compute and multi-vendor models.

Three Core Functional Modules

UModel: AI Inference Service Platform

Unifies the management of NVIDIA and domestic GPUs, supporting model import, one-click deployment, scaling, and monitoring to turn local models into standard API services.

UToken: Unified AI Token Gateway

Centralizes LLM requests from employees, systems, and Agents, providing unified authentication, smart routing, quota limits, billing, and security auditing.

O&M Monitoring: Observable Operations System

Centrally monitors core metrics like TTFT, TPOT, QPS, and GPU utilization, combining multi-dimensional alerts and log correlation to ensure stable service operations.

Use cases

End-to-end solutions from internal AI token governance to commercial compute monetization.

Unified Access for Internal AI Assistants

Unified Access for Internal AI Assistants

Employees, Office Assistants, Knowledge Bases, Coding Copilots

Single entry point for multiple applications, reducing repetitive integration costs.

01
Multi-Dept & Project Usage Governance
02
High-Compliance Private Deployment
03
Servitization of Enterprise Compute Pools
04
AI Token Commercial Sales
05
Channel / ISV / Agent Ecosystem Ops
06

From Compute Integration to Token Distribution: Rapidly Building an Enterprise AI Governance Platform

Combined with UModel, UToken helps enterprises build end-to-end workflows—from managing local GPUs, Private Cloud/UCloudStack, or cloud compute to model deployment and Token operations.

1

Integration

Integrates local GPUs, Private Cloud/UCloudStack, UCloud cloud GPUs, and AstraFlow, unifying scattered compute into a foundational resource pool.

2

Model Servitization

Utilizes UModel for one-click deployment, scaling, and monitoring, standardizing local model capabilities into accessible API services.

3

Unified Gateway Access

Employs UToken to provide smart routing, quota limits, and security guardrails, centrally managing requests from employees, systems, and Agents.

4

Ops Billing & Auditing

Tracks Token usage, billing, and audit logs comprehensively, ensuring AI invocations are measurable, auditable, and operable.

Not just deploying models, but building a deliverable, governable, and operable LLM inference infrastructure.

Five unified capabilities that turn scattered models and compute into enterprise-grade AI infrastructure.

Unified Supply

Unifies the management of open-source, proprietary, and industry models, publishing them as standardized services.

Unified Management

Integrates NVIDIA and domestic GPUs, reducing compute fragmentation and repetitive O&M.

Unified Governance

Centralizes authentication, routing, quotas, billing, and auditing, saving business teams from reinventing the wheel.

Security & Compliance

Supports local data retention, full-link auditing, permission isolation, and security guardrails to meet strict compliance needs.

Continuous Operations

Centrally monitors metrics like TTFT, TPOT, tok/s, QPS, and GPU utilization to ensure long-term stable operations.

$5绑卡奖励

下单前需绑定信用卡

绑定后即可使用支付宝、微信及东南亚本地钱包支付,并可联系客服领取 $5 绑卡奖励。

联系我们

Build an Enterprise AI Token Governance Platform, Starting with Unified Access

Whether for internal AI invocation governance or external AI Token commercial operations, UToken combined with UModel empowers enterprises to build end-to-end workflows—from compute management and model deployment to unified access, billing, and security auditing.