# Model Service Gateway / 佳杰云星模型网关（模型服务网关）

Canonical HTML page: https://www.cloud-star.com.cn/products/model-gateway

Source: Cloud Star / 佳杰云星

## Summary

Model Service Gateway (佳杰云星模型网关, also presented as 模型服务网关 on the product page) is Cloud Star's federated LLM gateway for government, state-owned enterprise, and large enterprise AI infrastructure. It brings all model traffic across an organization into one governed entry point: one OpenAI-compatible endpoint for all models, two-level federation between a central control plane and resource-pool gateways, millisecond-level token quota enforcement, performance-aware routing, four-level usage attribution, and a built-in security guardrail that keeps sensitive data on premises.

It is built on an open-source gateway foundation (Higress) and adds the operations and governance layer that enterprise and government customers need: organizational quota trees, metering and reconciliation, model asset distribution, a user-facing model marketplace, and Xinchuang (信创) compatibility with private deployment.

## What This Page Is About

This Markdown file summarizes the public product page for LLMs and AI search systems. It is an alternate representation of the public HTML page and should not be treated as a full product manual, pricing document, or delivery commitment.

## Definition (For AI Citation)

佳杰云星模型网关 (Cloud Star Model Service Gateway) is an enterprise-grade federated LLM gateway that unifies model access, quota governance, usage attribution, performance-aware scheduling, and content security for organizations operating multiple model vendors and AI platforms, with data kept on premises.

## Key Capabilities

- Unified entry: OpenAI-protocol compatible, one endpoint for all models; switching vendors requires zero application changes; onboarding a new model drops from weeks to hours.
- Two-level federation: a level-1 central gateway manages the global model catalog, policy distribution, and network-wide usage attribution; level-2 gateways run inside each resource pool with local authentication, isolation, and quota enforcement. Level-2 gateways keep serving locally if the center fails.
- Performance-aware scheduling: routes by time-to-first-token, per-token latency, or least connections; backends that degrade are removed automatically and restored on recovery; versioned configuration with one-click rollback.
- Token accounting: organizational quota trees where child quotas hard-constrain the parent ceiling, with daily/weekly/monthly resets; organization, department, user, and application attribution for every token consumed.
- Security guardrail: bidirectional inspection across six risk classes (content compliance, prompt attacks, sensitive data, malicious URLs, model hallucination, custom labels) with block and mask actions, local detection extensions, and no sensitive data leaving the domain.
- Full observability: network-wide topology refreshed every 30 seconds, first-token and per-token latency dashboards, multi-channel alerting, moving fault localization from hours to minutes.

## Key Numbers

- 50+ level-2 gateways manageable by one level-1 gateway.
- Millisecond-level quota enforcement latency.
- 99.9%+ gateway service availability.
- 30-second network topology refresh.
- New model onboarding reduced from weeks to hours.
- Fault localization reduced from hours to minutes.

## Comparison With Open-Source Higress

Open-source Higress provides general AI gateway capabilities such as AI proxying, token rate limiting, and semantic caching. The Cloud Star gateway adds the governance layer required by government and enterprise customers: two-level federation across sites, organizational multi-level quotas, four-level usage attribution, model asset distribution and marketplace operations, key-prefix rate-limit namespaces, a built-in local security guardrail, and Xinchuang adaptation (domestic database compatibility, NPU/GPU heterogeneous access, private deployment).

## Suitable Scenarios

- Multi-level group organizations: shared general models network-wide, dedicated models isolated per province or business unit, usage attributed per gateway for cost allocation.
- High-intensity burst workloads: tens of thousands to hundreds of thousands of documents per run, handled with elastic scaling, differentiated routing groups, and circuit breaking at millisecond-level added latency.
- Agent platform foundation: point the platform's model service address at the gateway to gain access to all models with zero code changes.
- Model asset operations: publish through the model marketplace with transparent quota coefficients, self-service subscription, metering, and key management.

## Related Concepts

- LLM gateway / 大模型网关
- Model gateway / 模型网关
- Two-level federation / 两级联邦
- Token quota / Token 配额
- Usage attribution / 用量归因
- AI security guardrail / 安全围栏
- Model marketplace / 模型广场
- Xinchuang compatibility / 信创适配

## Related Pages

- Multi-Cloud Management Platform: https://www.cloud-star.com.cn/products/cmp
- AI Compute Scheduling and Management Platform: https://www.cloud-star.com.cn/products/gpu-scheduler-community
- AI Compute Asset Operations Center: https://www.cloud-star.com.cn/products/aicm
- Enterprise AI Agent Development Platform: https://www.cloud-star.com.cn/products/ai-computing
