Feat : Gitbook

This commit is contained in:
decolua
2026-05-11 11:50:24 +07:00
parent 7ad538bcf2
commit fd92af77a0
124 changed files with 34154 additions and 4 deletions

View File

@@ -0,0 +1,537 @@
# Combos - Chuỗi Fallback Tùy chỉnh
Tạo các tổ hợp model tùy chỉnh với fallback tự động. Combo cho phép bạn định nghĩa chiến lược routing dựa trên chi phí, chất lượng và tính khả dụng.
---
## Combos là gì?
Combos là **chuỗi fallback tùy chỉnh** bạn tạo trong dashboard. Thay vì dùng một model duy nhất, bạn định nghĩa một chuỗi các model mà 9Router sẽ thử theo thứ tự.
**Ví dụ:**
```
Combo name: premium-coding
Models:
1. cc/claude-opus-4-5-20251101 (try first)
2. glm/glm-4.7 (if #1 quota exhausted)
3. minimax/MiniMax-M2.1 (if #2 quota exhausted)
```
**Dùng trong CLI:**
```
Model: premium-coding
```
9Router tự động thử từng model theo thứ tự cho đến khi thành công.
---
## Tại sao dùng Combos?
### 1. Tối đa hóa Giá trị Subscription
```
cc/claude-opus → glm/glm-4.7 → if/kimi-k2-thinking
→ Use subscription first, cheap backup, free emergency
→ Get full value from subscriptions you already pay for
```
### 2. Giảm Chi phí
```
glm/glm-4.7 → minimax/MiniMax-M2.1 → if/kimi-k2-thinking
→ Start with cheapest paid option ($0.60/1M)
→ Fallback to even cheaper ($0.20/1M)
→ Emergency free tier
→ Total cost: ~$5-10/month vs $2000 on ChatGPT API
```
### 3. Đảm bảo Khả dụng 24/7
```
cc/claude-opus → cx/gpt-5.2-codex → glm/glm-4.7 → if/kimi-k2-thinking
→ Always include free tier at the end
→ Never run out of quota
→ Code anytime, anywhere
```
### 4. Tối ưu Chất lượng
```
cc/claude-opus-4-5 → cx/gpt-5.2-codex → gc/gemini-3-pro
→ Best models first
→ Fallback to other premium models
→ Maintain high quality across fallback chain
```
---
## Cách tạo Combos
### Bước 1: Mở Dashboard
```
http://localhost:20128
→ Login with your password
```
### Bước 2: Đi đến Combos
```
Dashboard → Combos → Create New Combo
```
### Bước 3: Cấu hình Combo
**Tên Combo:**
```
premium-coding
```
**Mô tả (tùy chọn):**
```
Subscription first, cheap backup, free emergency
```
**Chọn Models:**
```
1. cc/claude-opus-4-5-20251101
2. glm/glm-4.7
3. minimax/MiniMax-M2.1
```
**Kéo để sắp xếp lại** - Ưu tiên từ trên xuống dưới.
### Bước 4: Lưu
```
Click "Save Combo"
→ Combo appears in model list
```
### Bước 5: Dùng trong CLI
```
Cursor/Cline/Any tool:
Model: premium-coding
```
---
## Ví dụ Combos
### Ví dụ 1: Premium Coding (Subscription → Cheap → Free)
**Mục tiêu**: Tối đa giá trị subscription, giảm chi phí thêm.
```
Dashboard → Combos → Create New
Name: premium-coding
Models:
1. cc/claude-opus-4-5-20251101
2. glm/glm-4.7
3. minimax/MiniMax-M2.1
```
**Sử dụng:**
```
Cursor IDE:
Model: premium-coding
```
**Hoạt động:**
```
Morning (fresh quota):
Request → cc/claude-opus-4-5 ✅
Afternoon (Claude quota out):
Request → glm/glm-4.7 ✅ (auto switched)
Evening (GLM quota out):
Request → minimax/MiniMax-M2.1 ✅ (auto switched)
```
**Chi phí hàng tháng (100M tokens):**
```
80M via Claude Code: $0 (subscription)
15M via GLM: $9
5M via MiniMax: $1
Total: $10 + your subscription
```
**Tiết kiệm**: ~99% so với ChatGPT API ($2000).
---
### Ví dụ 2: Budget Combo (Cheap → Free)
**Mục tiêu**: Giảm chi phí, dùng free tier làm backup.
```
Dashboard → Combos → Create New
Name: budget-combo
Models:
1. glm/glm-4.7
2. minimax/MiniMax-M2.1
3. if/kimi-k2-thinking
```
**Sử dụng:**
```
Cline:
Provider: OpenAI Compatible
Base URL: http://localhost:20128/v1
Model: budget-combo
```
**Hoạt động:**
```
Request → glm/glm-4.7
✅ Daily quota available → Use GLM ($0.60/1M)
❌ Quota exhausted → Try MiniMax ($0.20/1M)
❌ MiniMax quota out → Use iFlow (FREE)
```
**Chi phí hàng tháng (100M tokens):**
```
70M via GLM: $42
20M via MiniMax: $4
10M via iFlow: $0
Total: $46 vs $2000 on ChatGPT API
```
**Tiết kiệm**: 97%.
---
### Ví dụ 3: Free Combo (Chi phí 0)
**Mục tiêu**: 100% miễn phí, không bao giờ tốn tiền.
```
Dashboard → Combos → Create New
Name: free-combo
Models:
1. if/kimi-k2-thinking
2. qw/qwen3-coder-plus
3. kr/claude-sonnet-4.5
```
**Sử dụng:**
```
Claude Desktop:
Model: free-combo
```
**Hoạt động:**
```
Request → if/kimi-k2-thinking
✅ Available → Use iFlow
❌ Error → Try Qwen
❌ Error → Try Kiro
```
**Chi phí hàng tháng:**
```
100M tokens via free providers: $0
Total: $0 forever
```
**Use case**: Dự án cá nhân, học tập, thử nghiệm.
---
### Ví dụ 4: Quality First (Chỉ Premium Models)
**Mục tiêu**: Chất lượng tốt nhất, không fallback rẻ.
```
Dashboard → Combos → Create New
Name: quality-first
Models:
1. cc/claude-opus-4-5-20251101
2. cx/gpt-5.2-codex
3. gc/gemini-3-pro-preview
```
**Sử dụng:**
```
Codex CLI:
export OPENAI_BASE_URL="http://localhost:20128"
Model: quality-first
```
**Hoạt động:**
```
Request → cc/claude-opus-4-5
❌ Quota out → cx/gpt-5.2-codex
❌ Quota out → gc/gemini-3-pro-preview
❌ All out → Return error (no cheap fallback)
```
**Use case**: Code production quan trọng, refactoring phức tạp.
---
### Ví dụ 5: Multi-Subscription (Tối đa hết tất cả)
**Mục tiêu**: Dùng hết subscription trước khi trả thêm tiền.
```
Dashboard → Combos → Create New
Name: multi-sub
Models:
1. gc/gemini-3-flash-preview (FREE 180K/month)
2. cc/claude-opus-4-5-20251101 (Pro subscription)
3. cx/gpt-5.2-codex (Plus subscription)
4. gh/gpt-5 (Copilot subscription)
5. glm/glm-4.7 (Cheap backup)
6. if/kimi-k2-thinking (Free emergency)
```
**Chi phí hàng tháng (200M tokens):**
```
50M via Gemini CLI: $0 (free tier)
80M via Claude Code: $0 (subscription)
40M via Codex: $0 (subscription)
20M via Copilot: $0 (subscription)
8M via GLM: $4.80
2M via iFlow: $0
Total: $4.80 + existing subscriptions
```
**Kết quả**: Dùng 190M tokens từ subscription, chỉ $4.80 phụ.
---
### Ví dụ 6: Tối ưu Reset Quota
**Mục tiêu**: Phân bổ sử dụng dựa trên thời gian reset.
```
Dashboard → Combos → Create New
Name: reset-optimized
Models:
1. cc/claude-opus-4-5 (5h reset, use morning)
2. gc/gemini-3-flash (1K/day, use afternoon)
3. glm/glm-4.7 (daily 10AM reset, use evening)
4. minimax/MiniMax-M2.1 (5h rolling, use night)
5. if/kimi-k2-thinking (unlimited, emergency)
```
**Lịch trình hàng ngày:**
```
08:00 - 13:00: Claude Code (fresh 5h quota)
13:00 - 18:00: Gemini CLI (1K/day quota)
18:00 - 22:00: GLM (resets 10AM next day)
22:00 - 08:00: MiniMax (5h rolling) or iFlow
```
**Kết quả**: Code 24/7 với chi phí tối thiểu.
---
## Dùng Combos trong CLI Tools
### Cursor IDE
```
Settings → Models → Advanced:
OpenAI API Base URL: http://localhost:20128/v1
OpenAI API Key: [from dashboard]
Model: premium-coding
```
### Claude Desktop
Sửa `~/.claude/config.json`:
```json
{
"anthropic_api_base": "http://localhost:20128/v1",
"anthropic_api_key": "your-9router-api-key",
"model": "budget-combo"
}
```
### Codex CLI
```bash
export OPENAI_BASE_URL="http://localhost:20128"
export OPENAI_API_KEY="your-9router-api-key"
codex --model quality-first "your prompt"
```
### Cline / Continue / RooCode
```
Provider: OpenAI Compatible
Base URL: http://localhost:20128/v1
API Key: [from dashboard]
Model: free-combo
```
### API Request
```bash
curl http://localhost:20128/v1/chat/completions \
-H "Authorization: Bearer your-api-key" \
-H "Content-Type: application/json" \
-d '{
"model": "premium-coding",
"messages": [
{"role": "user", "content": "Write a function to..."}
],
"stream": true
}'
```
---
## Best Practices
### 1. Luôn bao gồm Free Tier
```
✅ Good:
cc/claude-opus → glm/glm-4.7 → if/kimi-k2-thinking
❌ Bad:
cc/claude-opus → glm/glm-4.7
(no free fallback, can run out of quota)
```
**Lý do**: Đảm bảo khả dụng 24/7, không bao giờ bị chặn bởi quota.
### 2. Sắp xếp theo Chi phí (Rẻ đến Đắt)
```
✅ Good:
glm/glm-4.7 → minimax/MiniMax-M2.1 → cc/claude-opus
❌ Bad:
cc/claude-opus → glm/glm-4.7
(wastes subscription quota on simple tasks)
```
**Ngoại lệ**: Nếu muốn tối đa giá trị subscription, đặt subscription đầu tiên.
### 3. Phù hợp với Yêu cầu Chất lượng
```
For production code:
cc/claude-opus → cx/gpt-5.2-codex → glm/glm-4.7
For quick tasks:
glm/glm-4.7 → if/kimi-k2-thinking
For experimentation:
if/kimi-k2-thinking → qw/qwen3-coder-plus
```
### 4. Cân nhắc Thời gian Reset Quota
```
Morning combo (fresh quotas):
cc/claude-opus → cx/gpt-5.2-codex
Evening combo (quotas likely exhausted):
glm/glm-4.7 → minimax/MiniMax-M2.1 → if/kimi-k2-thinking
```
### 5. Tạo nhiều Combo cho các Use Case khác nhau
```
premium-coding: For complex tasks
budget-combo: For simple tasks
free-combo: For experimentation
quality-first: For production code
```
**Chuyển đổi combo** dựa trên yêu cầu task.
### 6. Theo dõi hiệu năng Combo
```
Dashboard → Analytics → Combo Usage:
premium-coding:
80% via cc/claude-opus (good, using subscription)
15% via glm/glm-4.7 (acceptable backup)
5% via minimax (rare fallback)
```
**Tối ưu**: Nếu fallback quá nhiều, tăng quota chính hoặc sắp xếp lại model.
---
## Cấu hình Nâng cao
### Đặt Giới hạn Ngân sách cho mỗi Combo
```
Dashboard → Combos → Edit → Budget:
Daily limit: $5
Monthly limit: $50
```
Khi đạt giới hạn, 9Router bỏ qua model trả phí và chỉ dùng free tier.
### Bật/Tắt Model trong Combo
```
Dashboard → Combos → Edit → Models:
✅ cc/claude-opus-4-5 (enabled)
❌ glm/glm-4.7 (temporarily disabled)
✅ if/kimi-k2-thinking (enabled)
```
**Use case**: Tạm tắt model đắt mà không cần xóa combo.
### Clone Combo có sẵn
```
Dashboard → Combos → Clone "premium-coding"
→ Creates copy with "-copy" suffix
→ Modify and save as new combo
```
**Use case**: Tạo biến thể cho các kịch bản khác nhau.
---
## Troubleshooting
**Issue: Combo không xuất hiện trong danh sách model**
**Giải pháp:**
1. Refresh dashboard
2. Kiểm tra combo đã được lưu (dấu tick xanh)
3. Khởi động lại CLI tool để refresh danh sách model
**Issue: Combo luôn dùng model cuối cùng (free tier)**
**Giải pháp:**
1. Kiểm tra quota cho các model chính (Dashboard → Quota)
2. Xác minh API keys hợp lệ (Dashboard → Providers)
3. Kiểm tra giới hạn ngân sách không vượt quá
**Issue: Combo tốn hơn dự kiến**
**Giải pháp:**
1. Dashboard → Analytics → Xem usage combo
2. Kiểm tra model chính có bị hết quota không
3. Sắp xếp lại model (đặt rẻ hơn lên trước)
4. Đặt giới hạn ngân sách
---
## Liên quan
- [Smart Routing](./smart-routing.md) - Cách auto fallback hoạt động
- [Quota Tracking](./quota-tracking.md) - Theo dõi sử dụng và chi phí

View File

@@ -0,0 +1,687 @@
# Quota Tracking & Giám sát Usage
Theo dõi tiêu thụ token thời gian thực, giám sát giới hạn quota, ước tính chi phí và nhận cảnh báo trước khi hết. Không bao giờ lãng phí quota subscription hoặc vượt giới hạn ngân sách.
---
## Tổng quan
9Router cung cấp quota tracking toàn diện cho mọi provider:
- **Tiêu thụ token thời gian thực** - Xem tokens dùng mỗi request
- **Giới hạn quota & còn lại** - Theo dõi usage so với giới hạn
- **Đếm ngược Reset** - Biết khi nào quota refresh
- **Ước tính chi phí** - Tính chi tiêu cho tier trả phí
- **Báo cáo hàng tháng** - Phân tích pattern sử dụng
- **Cảnh báo & thông báo** - Nhận cảnh báo trước giới hạn
---
## Tổng quan Dashboard
### Tóm tắt Quota
```
Dashboard → Home → Quota Overview
┌─────────────────────────────────────────────┐
│ Claude Code (cc/) │
│ ████████████░░░░░░░░ 2.5h / 5h (50%) │
│ Resets in: 2h 30m │
│ Cost: $0 (subscription) │
└─────────────────────────────────────────────┘
┌─────────────────────────────────────────────┐
│ Gemini CLI (gc/) │
│ ████████░░░░░░░░░░░░ 450 / 1000 (45%) │
│ Daily reset in: 18h 30m │
│ Monthly: 45K / 180K (25%) │
│ Cost: $0 (free tier) │
└─────────────────────────────────────────────┘
┌─────────────────────────────────────────────┐
│ GLM-4.7 (glm/) │
│ ██████████████░░░░░░ 7M / 10M tokens (70%) │
│ Resets: Daily 10:00 AM (in 5h 35m) │
│ Cost today: $4.20 │
└─────────────────────────────────────────────┘
┌─────────────────────────────────────────────┐
│ MiniMax M2.1 (minimax/) │
│ ████████████████░░░░ 4M / 5M tokens (80%) │
│ Rolling 5h window │
│ Cost (5h): $0.80 │
└─────────────────────────────────────────────┘
┌─────────────────────────────────────────────┐
│ iFlow (if/) │
│ ████████████████████ Unlimited │
│ Cost: $0 (free forever) │
└─────────────────────────────────────────────┘
```
---
## Tiêu thụ Token Thời gian thực
### Theo dõi từng Request
Mỗi request hiển thị usage token chi tiết:
```
Dashboard → Activity → Recent Requests
Request #1234
Model: cc/claude-opus-4-5-20251101
Timestamp: 2026-02-04 04:15:32
Tokens:
Input: 1,250 tokens
Output: 850 tokens
Total: 2,100 tokens
Cost: $0 (subscription quota)
Duration: 3.2s
Status: ✅ Success
```
### Live Usage Monitor
```
Dashboard → Live Monitor
Current request:
Model: glm/glm-4.7
Tokens streamed: 450 / ~800 estimated
Cost so far: $0.0009
Duration: 1.8s
```
### Phân tích Token theo Model
```
Dashboard → Analytics → Token Usage
Today (Feb 4, 2026):
cc/claude-opus-4-5: 15M tokens ($0, subscription)
glm/glm-4.7: 8M tokens ($4.80)
if/kimi-k2-thinking: 3M tokens ($0, free)
Total: 26M tokens
Cost: $4.80
```
---
## Giới hạn Quota & Thời gian Reset
### Subscription Providers
**Claude Code (Pro/Max)**
```
Quota type: Time-based (5-hour rolling)
Limit: 5 hours of usage
Reset: Rolling 5-hour window + Weekly refresh
Tracking: Usage time per model
Dashboard shows:
Opus: 2.5h / 5h used
Sonnet: 1.2h / 5h used
Haiku: 0.8h / 5h used
Weekly reset: Every Monday 00:00 UTC
```
**OpenAI Codex (Plus/Pro)**
```
Quota type: Time-based (5-hour rolling)
Limit: 5 hours (Plus) / 10 hours (Pro)
Reset: Rolling 5-hour window + Weekly refresh
Dashboard shows:
GPT-5.2 Codex: 3.5h / 5h used
Resets in: 1h 30m
```
**Gemini CLI (MIỄN PHÍ)**
```
Quota type: Request count + Monthly tokens
Daily limit: 1,000 requests
Monthly limit: 180,000 completions
Reset: Daily 00:00 UTC + Monthly 1st
Dashboard shows:
Today: 450 / 1,000 requests (45%)
This month: 45K / 180K completions (25%)
Daily reset in: 18h 30m
Monthly reset in: 26 days
```
**GitHub Copilot**
```
Quota type: Monthly usage
Limit: Varies by plan
Reset: 1st of each month
Dashboard shows:
Usage: 60% of monthly quota
Resets: March 1, 2026 (in 25 days)
```
### Cheap Providers
**GLM-4.7**
```
Quota type: Daily token limit
Limit: 10M tokens/day (Coding Plan)
Reset: Daily 10:00 AM Beijing Time (UTC+8)
Dashboard shows:
Used: 7M / 10M tokens (70%)
Remaining: 3M tokens
Resets in: 5h 35m
Cost today: $4.20
```
**MiniMax M2.1**
```
Quota type: Rolling 5-hour window
Limit: 5M tokens per 5 hours
Reset: Continuous rolling window
Dashboard shows:
Used (5h): 4M / 5M tokens (80%)
Oldest usage expires in: 45m
Cost (5h): $0.80
```
**Kimi K2**
```
Quota type: Monthly subscription
Limit: 10M tokens/month ($9 flat)
Reset: Monthly on subscription date
Dashboard shows:
Used: 6M / 10M tokens (60%)
Resets: Feb 15, 2026 (in 11 days)
Cost: $9/month (prepaid)
```
### Free Providers
**iFlow / Qwen / Kiro**
```
Quota type: Unlimited (rate-limited)
Limit: No hard limit
Reset: N/A
Dashboard shows:
Used today: 5M tokens
Cost: $0 (free forever)
Status: ✅ Available
```
---
## Ước tính Chi phí
### Theo dõi Chi phí Thời gian thực
```
Dashboard → Costs → Today
Subscription providers: $0
Claude Code: 15M tokens ($0, included)
Gemini CLI: 3M tokens ($0, free tier)
Paid providers: $4.80
GLM-4.7: 8M tokens ($4.80)
Input: 6M × $0.60/1M = $3.60
Output: 2M × $2.20/1M = $4.40
Total: $4.80
Free providers: $0
iFlow: 3M tokens ($0)
Total today: $4.80
```
### Báo cáo Chi tiêu Hàng tháng
```
Dashboard → Costs → This Month (February 2026)
Week 1 (Feb 1-7):
Subscription: $0 (80M tokens)
Paid: $15.20 (25M tokens)
Free: $0 (10M tokens)
Total: $15.20
Week 2 (Feb 8-14):
Subscription: $0 (75M tokens)
Paid: $12.80 (20M tokens)
Free: $0 (8M tokens)
Total: $12.80
Month to date: $28.00
Projected (30 days): ~$120
Breakdown by provider:
GLM-4.7: $22.00 (78%)
MiniMax M2.1: $6.00 (22%)
Average cost per 1M tokens: $0.62
Savings vs ChatGPT API: 97% ($4,000 → $120)
```
### Dự kiến Chi phí
```
Dashboard → Costs → Projections
Based on last 7 days usage:
Daily average: 50M tokens
Daily cost: $4.50
Monthly projection:
Tokens: 1,500M (1.5B)
Cost: $135
Breakdown:
Subscription: 900M tokens ($0)
GLM-4.7: 450M tokens ($90)
MiniMax: 120M tokens ($24)
Free: 30M tokens ($0)
Budget status:
Daily limit: $5 → 90% used today
Monthly limit: $150 → 90% projected
⚠️ Warning: May exceed monthly budget
```
---
## Dashboard Usage
### Thống kê Tổng quan
```
Dashboard → Analytics → Overview
Today (Feb 4, 2026):
Requests: 1,234
Tokens: 26M
Cost: $4.80
Avg response time: 2.1s
This week:
Requests: 8,456
Tokens: 180M
Cost: $28.00
Success rate: 99.2%
This month:
Requests: 15,234
Tokens: 320M
Cost: $52.00
Top model: cc/claude-opus-4-5 (45%)
```
### Usage theo Model
```
Dashboard → Analytics → Models
Top models (this month):
1. cc/claude-opus-4-5: 145M tokens (45%)
2. glm/glm-4.7: 95M tokens (30%)
3. if/kimi-k2-thinking: 50M tokens (16%)
4. minimax/MiniMax-M2.1: 20M tokens (6%)
5. gc/gemini-3-flash: 10M tokens (3%)
Cost breakdown:
cc/claude-opus: $0 (subscription)
glm/glm-4.7: $45.00
if/kimi-k2-thinking: $0 (free)
minimax/MiniMax-M2.1: $7.00
gc/gemini-3-flash: $0 (free)
```
### Usage theo Thời gian
```
Dashboard → Analytics → Timeline
Hourly usage (today):
00:00 - 01:00: 0.5M tokens
01:00 - 02:00: 0.2M tokens
...
08:00 - 09:00: 3.2M tokens (peak)
09:00 - 10:00: 2.8M tokens
...
23:00 - 00:00: 0.8M tokens
Peak hours: 08:00 - 12:00 (morning coding)
Low hours: 00:00 - 06:00 (night)
```
### Usage theo Combo
```
Dashboard → Analytics → Combos
premium-coding:
Requests: 456
Tokens: 12M
Cost: $2.40
Breakdown:
cc/claude-opus: 8M tokens (67%, $0)
glm/glm-4.7: 3M tokens (25%, $1.80)
minimax/MiniMax-M2.1: 1M tokens (8%, $0.20)
budget-combo:
Requests: 234
Tokens: 6M
Cost: $1.20
Breakdown:
glm/glm-4.7: 4M tokens (67%, $2.40)
if/kimi-k2-thinking: 2M tokens (33%, $0)
```
---
## Cảnh báo & Thông báo
### Cảnh báo Quota
```
Dashboard → Settings → Alerts
Quota warnings:
✅ Alert at 80% quota used
✅ Alert at 90% quota used
✅ Alert when quota exhausted
✅ Notify when quota resets
Delivery:
✅ Dashboard notification
✅ Email (optional)
✅ Webhook (optional)
```
**Ví dụ thông báo:**
```
⚠️ Claude Code quota 80% used
2.5h remaining (resets in 1h 30m)
⚠️ GLM-4.7 quota 90% used
1M tokens remaining (resets in 5h)
✅ Gemini CLI quota reset
1,000 requests available (daily limit)
```
### Cảnh báo Ngân sách
```
Dashboard → Settings → Budget Alerts
Daily budget: $5
✅ Alert at 80% ($4)
✅ Alert at 100% ($5)
✅ Auto-switch to free tier when exceeded
Monthly budget: $150
✅ Alert at 50% ($75)
✅ Alert at 80% ($120)
✅ Alert at 100% ($150)
```
**Ví dụ thông báo:**
```
⚠️ Daily budget 80% used
$4.00 / $5.00 spent today
⚠️ Monthly budget 50% reached
$75 / $150 spent this month
Projected: $135 (within budget)
🚨 Daily budget exceeded
$5.20 / $5.00 spent today
Auto-switched to free tier
```
### Phát hiện Bất thường Chi phí
```
Dashboard → Settings → Anomaly Detection
✅ Detect unusual spending patterns
✅ Alert on cost spikes (>2× daily average)
✅ Warn on quota exhaustion patterns
Example alert:
⚠️ Cost spike detected
Today: $12.50 (2.5× daily average)
Reason: High GLM-4.7 usage (20M tokens)
Suggestion: Check if primary models quota-exhausted
```
---
## Best Practices
### 1. Theo dõi Quota Hàng ngày
```
Daily routine:
1. Check dashboard quota overview (30 seconds)
2. Review reset times
3. Plan usage around quota availability
```
**Ví dụ:**
```
Morning check:
✅ Claude Code: 5h available (fresh reset)
✅ Gemini CLI: 1K requests available
⚠️ GLM-4.7: 2M tokens left (resets 10AM)
Action: Use Claude Code for morning work
```
### 2. Đặt Giới hạn Ngân sách
```
Dashboard → Settings → Budget:
Daily: $5 (prevents overspending)
Monthly: $150 (aligns with budget)
```
**Kết quả**: Auto-switch sang free tier khi đạt giới hạn.
### 3. Tối ưu Combo Usage
```
Dashboard → Analytics → Combos:
Review which models are used most
Adjust combo order to minimize costs
```
**Ví dụ:**
```
Current: cc/claude-opus → glm/glm-4.7
80% via Claude (good)
20% via GLM ($12/month)
Optimized: gc/gemini-3-flash → cc/claude-opus → glm/glm-4.7
50% via Gemini (free)
40% via Claude (subscription)
10% via GLM ($6/month)
Savings: $6/month
```
### 4. Theo dõi Thời gian Reset
```
Dashboard → Quota → Reset Schedule:
Claude Code: 5h rolling + Weekly Monday
Gemini CLI: Daily 00:00 UTC + Monthly 1st
GLM-4.7: Daily 10:00 AM Beijing Time
MiniMax: Rolling 5h window
```
**Chiến lược**: Dùng provider khi quota mới reset.
### 5. Xem Báo cáo Hàng tháng
```
Dashboard → Analytics → Monthly Report:
Total tokens: 1.5B
Total cost: $120
Savings: 97% vs ChatGPT API
Insights:
- 60% usage via subscriptions ($0)
- 30% via GLM ($90)
- 10% via free tier ($0)
Optimization:
- Increase Gemini CLI usage (free)
- Reduce GLM usage (expensive)
```
---
## Truy cập API
### Lấy trạng thái Quota
```bash
GET http://localhost:20128/api/quota
Authorization: Bearer your-api-key
Response:
{
"providers": [
{
"id": "cc",
"name": "Claude Code",
"quota": {
"used": 2.5,
"limit": 5,
"unit": "hours",
"percentage": 50
},
"reset": {
"type": "rolling",
"window": "5h",
"nextReset": "2026-02-04T06:45:00Z"
},
"cost": {
"today": 0,
"month": 0,
"currency": "USD"
}
},
{
"id": "glm",
"name": "GLM-4.7",
"quota": {
"used": 7000000,
"limit": 10000000,
"unit": "tokens",
"percentage": 70
},
"reset": {
"type": "daily",
"time": "10:00 AM UTC+8",
"nextReset": "2026-02-04T10:00:00+08:00"
},
"cost": {
"today": 4.20,
"month": 52.00,
"currency": "USD"
}
}
]
}
```
### Lấy Usage Stats
```bash
GET http://localhost:20128/api/usage?period=today
Authorization: Bearer your-api-key
Response:
{
"period": "today",
"date": "2026-02-04",
"summary": {
"requests": 1234,
"tokens": 26000000,
"cost": 4.80
},
"byModel": [
{
"model": "cc/claude-opus-4-5",
"requests": 456,
"tokens": 15000000,
"cost": 0
},
{
"model": "glm/glm-4.7",
"requests": 234,
"tokens": 8000000,
"cost": 4.80
}
]
}
```
---
## Troubleshooting
**Issue: Quota hiển thị 0% nhưng request thất bại**
**Giải pháp:**
1. Kiểm tra kết nối provider (Dashboard → Providers)
2. Xác minh API keys hợp lệ
3. Kiểm tra provider có down không (trang status)
4. Thử kết nối lại OAuth providers
**Issue: Ước tính chi phí sai**
**Giải pháp:**
1. Dashboard → Settings → Pricing
2. Xác minh giá mỗi provider khớp với mức hiện tại
3. Cập nhật giá nếu provider thay đổi
4. Liên hệ support nếu vẫn lệch
**Issue: Thời gian reset không cập nhật**
**Giải pháp:**
1. Refresh dashboard (F5)
2. Kiểm tra thời gian hệ thống đúng
3. Xác minh cài đặt timezone
4. Khởi động lại 9Router nếu vẫn lỗi
**Issue: Không nhận được cảnh báo**
**Giải pháp:**
1. Dashboard → Settings → Alerts
2. Xác minh địa chỉ email đúng
3. Kiểm tra folder spam
4. Test notification (nút Send Test)
---
## Liên quan
- [Smart Routing](./smart-routing.md) - Auto fallback dựa trên quota
- [Combos](./combos.md) - Tạo chuỗi fallback tùy chỉnh

View File

@@ -0,0 +1,407 @@
# Smart Routing & Auto Fallback
9Router tự động định tuyến request qua provider tốt nhất hiện có bằng hệ thống fallback 3 tầng. Không bao giờ ngừng code vì giới hạn quota hay rate limiting.
---
## Cách hoạt động
9Router dùng định tuyến thông minh để tối đa hóa subscription hiện có, giảm chi phí và đảm bảo khả dụng 24/7:
```
Request → 9Router → Check Tier 1 (Subscription)
↓ quota exhausted
Check Tier 2 (Cheap)
↓ budget limit
Check Tier 3 (Free)
↓
Response
```
### Hệ thống Fallback 3 tầng
**Tier 1: SUBSCRIPTION (Chính)**
- Claude Code (Pro/Max)
- OpenAI Codex (Plus/Pro)
- Gemini CLI (MIỄN PHÍ 180K/tháng)
- GitHub Copilot
- Antigravity (Google)
**Mục tiêu**: Tối đa giá trị từ subscription đã trả tiền.
**Tier 2: CHEAP (Backup)**
- GLM-4.7 ($0.60/1M input)
- MiniMax M2.1 ($0.20/1M input)
- Kimi K2 ($9/tháng cố định)
**Mục tiêu**: Backup siêu rẻ khi hết quota subscription (~90% rẻ hơn ChatGPT API).
**Tier 3: FREE (Khẩn cấp)**
- iFlow (8 models)
- Qwen (3 models)
- Kiro (Claude MIỄN PHÍ)
**Mục tiêu**: Fallback chi phí 0 để code không giới hạn.
---
## Chuyển đổi Tự động
9Router giám sát quota thời gian thực và chuyển provider tự động:
### Kịch bản 1: Hết Quota Subscription
```
User request → cc/claude-opus-4-5
↓ quota exhausted (5-hour limit reached)
Auto switch → glm/glm-4.7
↓ daily quota exhausted
Auto switch → minimax/MiniMax-M2.1
↓ 5-hour quota exhausted
Auto switch → if/kimi-k2-thinking (FREE)
↓
Response delivered ✅
```
**Kết quả**: Zero downtime, trải nghiệm liền mạch.
### Kịch bản 2: Rate Limiting
```
User request → cx/gpt-5.2-codex
↓ rate limited (too many requests)
Auto switch → glm/glm-4.7
↓
Response delivered ✅
```
### Kịch bản 3: Provider không khả dụng
```
User request → cc/claude-opus-4-5
↓ provider error (503)
Auto switch → next available model
↓
Response delivered ✅
```
---
## Logic chọn Model
9Router chọn model tốt nhất dựa trên:
1. **Khả dụng quota** - Kiểm tra provider còn quota không
2. **Tier chi phí** - Ưu tiên subscription → cheap → free
3. **Thời gian reset** - Cân nhắc khi quota reset
4. **Sức khỏe provider** - Bỏ qua provider có lỗi
### Ví dụ Thứ tự Ưu tiên
Cho request đến `cc/claude-opus-4-5`:
```
1. Check Claude Code quota
✅ Available → Use cc/claude-opus-4-5
❌ Exhausted → Continue to step 2
2. Check fallback tier (if configured)
✅ GLM quota available → Use glm/glm-4.7
❌ Exhausted → Continue to step 3
3. Check free tier
✅ iFlow available → Use if/kimi-k2-thinking
❌ All exhausted → Return quota error
```
---
## Tùy chọn Cấu hình
### Cài đặt Dashboard
**1. Bật/Tắt Auto Fallback**
```
Dashboard → Settings → Smart Routing
→ Toggle "Auto Fallback" ON/OFF
```
- **ON** (mặc định): Chuyển tier tự động
- **OFF**: Strict mode, trả lỗi nếu model chính không khả dụng
**2. Đặt Giới hạn Ngân sách**
```
Dashboard → Settings → Budget Control
→ Daily limit: $5
→ Monthly limit: $50
```
Khi đạt ngân sách, 9Router tự động chuyển sang free tier.
**3. Cấu hình Thứ tự Fallback**
```
Dashboard → Settings → Fallback Priority
→ Drag to reorder providers within each tier
```
Ví dụ thứ tự tùy chỉnh:
```
Tier 1: Gemini CLI → Claude Code → Codex
Tier 2: MiniMax → GLM → Kimi
Tier 3: iFlow → Kiro → Qwen
```
**4. Thông báo Reset Quota**
```
Dashboard → Settings → Notifications
→ Email when quota resets
→ Alert when 80% quota used
```
---
## Ví dụ
### Ví dụ 1: Auto Fallback Cơ bản
**Setup:**
```
Model: cc/claude-opus-4-5-20251101
Fallback: Auto (default 3-tier)
```
**Hoạt động:**
```
Morning (fresh quota):
Request → cc/claude-opus-4-5 ✅
Afternoon (quota exhausted):
Request → glm/glm-4.7 ✅ (auto switched)
Evening (GLM quota out):
Request → minimax/MiniMax-M2.1 ✅ (auto switched)
Late night (all paid quota out):
Request → if/kimi-k2-thinking ✅ (free tier)
```
**Chi phí**: ~$5-10/tháng extra (chủ yếu được bao bởi subscription).
### Ví dụ 2: Định tuyến theo Ngân sách
**Setup:**
```
Dashboard → Settings:
Daily budget: $2
Monthly budget: $20
Fallback: Enabled
```
**Hoạt động:**
```
Day 1-15 (within budget):
Requests → glm/glm-4.7 (cheap tier)
Cost: $1.50/day
Day 16 (budget reached):
Requests → if/kimi-k2-thinking (free tier)
Cost: $0
Next month (budget resets):
Requests → glm/glm-4.7 again
```
**Kết quả**: Không bao giờ vượt $20/tháng, luôn khả dụng.
### Ví dụ 3: Chế độ Chỉ Subscription
**Setup:**
```
Dashboard → Settings:
Auto Fallback: OFF
Strict mode: ON
```
**Hoạt động:**
```
Request → cc/claude-opus-4-5
✅ Quota available → Success
❌ Quota exhausted → Return error (no fallback)
```
**Use case**: Khi chỉ muốn dùng subscription trả phí, không phí thêm.
### Ví dụ 4: Chế độ Chỉ Free
**Setup:**
```
Model: if/kimi-k2-thinking
Fallback: qw/qwen3-coder-plus → kr/claude-sonnet-4.5
```
**Hoạt động:**
```
All requests → Free tier only
Cost: $0 forever
```
**Use case**: Dự án cá nhân, học tập, thử nghiệm.
---
## Best Practices
### 1. Tối đa Giá trị Subscription
```
Strategy:
- Set subscription models as Tier 1
- Monitor quota usage in dashboard
- Use cheap tier only when subscription exhausted
```
**Ví dụ combo:**
```
cc/claude-opus-4-5 → glm/glm-4.7 → if/kimi-k2-thinking
```
### 2. Tối ưu Chi phí
```
Strategy:
- Use Gemini CLI free tier first (180K/month)
- Fallback to GLM/MiniMax (ultra-cheap)
- Emergency: iFlow (free)
```
**Ví dụ combo:**
```
gc/gemini-3-flash-preview → glm/glm-4.7 → if/kimi-k2-thinking
```
### 3. Tối ưu Chất lượng
```
Strategy:
- Use best models (Claude Opus, GPT-5.2)
- Fallback to good cheap models (GLM-4.7)
- Last resort: Free tier
```
**Ví dụ combo:**
```
cc/claude-opus-4-5 → cx/gpt-5.2-codex → glm/glm-4.7
```
### 4. Khả dụng 24/7
```
Strategy:
- Always include free tier in fallback
- Monitor quota reset times
- Distribute usage across providers
```
**Ví dụ combo:**
```
cc/claude-opus-4-5 → glm/glm-4.7 → minimax/MiniMax-M2.1 → if/kimi-k2-thinking
```
**Kết quả**: Không bao giờ hết quota, code mọi lúc.
---
## Chiến lược Reset Quota
Lên kế hoạch usage quanh thời gian reset quota:
| Provider | Quota Reset | Chiến lược |
|----------|-------------|----------|
| **Claude Code** | 5 giờ + hàng tuần | Dùng buổi sáng, quota mới |
| **Codex** | 5 giờ + hàng tuần | Dùng sau khi hết quota Claude |
| **Gemini CLI** | Hàng ngày (1K) + Hàng tháng (180K) | Dùng cả ngày |
| **GLM-4.7** | Hàng ngày 10:00 AM | Dùng buổi tối, reset sáng hôm sau |
| **MiniMax M2.1** | 5 giờ rolling | Dùng mọi lúc, theo rolling window |
| **iFlow/Qwen/Kiro** | Không giới hạn | Backup khẩn cấp |
**Ví dụ lịch hàng ngày:**
```
08:00 - 13:00: Claude Code (fresh 5h quota)
13:00 - 18:00: Gemini CLI (1K/day quota)
18:00 - 22:00: GLM-4.7 (cheap, resets 10AM)
22:00 - 08:00: MiniMax or iFlow (5h rolling or free)
```
---
## Giám sát & Cảnh báo
### Dashboard Quota Tracker
```
Dashboard → Quota Overview:
Claude Code: 2.5h / 5h remaining (50%)
Gemini CLI: 450 / 1000 requests today
GLM-4.7: 5M / 10M tokens (resets in 8h)
MiniMax: 3M / 5M tokens (rolling 5h)
```
### Thông báo Thời gian thực
```
Dashboard → Notifications:
⚠️ Claude Code quota 80% used (1h remaining)
✅ GLM-4.7 quota reset (10M tokens available)
💰 Daily budget 50% used ($2.50 / $5)
```
### Usage Analytics
```
Dashboard → Analytics:
Today: 50M tokens
- 30M via Claude Code (subscription)
- 15M via GLM-4.7 ($9)
- 5M via iFlow (free)
Cost: $9 (vs $1000 on ChatGPT API)
Savings: 99%
```
---
## Troubleshooting
**Issue: "All providers quota exhausted"**
**Giải pháp:**
1. Kiểm tra quota tracker trong dashboard
2. Đợi quota reset (xem countdown)
3. Thêm free tier vào fallback chain
4. Hoặc tăng giới hạn ngân sách
**Issue: "Too many fallback switches"**
**Giải pháp:**
1. Kiểm tra provider chính có down không
2. Tăng giới hạn quota (upgrade subscription)
3. Dùng model chính rẻ hơn (GLM thay vì Claude)
**Issue: "Unexpected costs"**
**Giải pháp:**
1. Dashboard → Analytics → Xem usage
2. Đặt giới hạn ngân sách hàng ngày/tháng
3. Chuyển sang free tier cho task không quan trọng
4. Dùng combo với free fallback
---
## Liên quan
- [Combos](./combos.md) - Tạo chuỗi fallback tùy chỉnh
- [Quota Tracking](./quota-tracking.md) - Theo dõi usage và chi phí