diff --git a/CHANGELOG.md b/CHANGELOG.md index e7d4bea..e7cd72f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,6 +6,12 @@ --- +## v0.8.0 (2026-10-03) + +- **模型映射新增「最大输出 Token」列**:逐模型设置 `max_output_size`(写入 config.toml,留空用上游默认),models.dev 有 output 上限时显示可点击参考值一键填入;获取模型列表时自动预填 +- **models.dev 快照新增 output 上限字段**:同步至 2026-10-03(8,385 模型 / 226 供应商,output 覆盖 97.4%);应用内同步(Rust 侧)同步支持,避免同步后字段丢失 +- **移除无效的「声明支持 1M」复选框**:该控件读时由上下文长度派生、写时不持久化,勾选从不生效,由「最大输出 Token」列取代 + ## v0.7.23 (2026-09-30) - **思考等级五档 + 能力感知**:全局配置的思考等级扩展为低 / 中 / 高 / Max / XHigh,按当前默认模型声明的 `support_efforts` 自动置灰不支持档位并提示原因;models.dev 标记不支持思考的模型整体禁用思考区;未声明档位的模型保留可选并提示"上游可能拒绝" diff --git a/README.md b/README.md index d9360d9..700cb33 100644 --- a/README.md +++ b/README.md @@ -79,7 +79,7 @@ | **连通性测试** | 实测 `base_url` 延迟,绿/橙/红彩色气泡,6 秒自动消失 | | **重复供应商** | 一键深拷贝供应商 + 全部模型,key 改 `xxx-copy` | | **图标按钮操作** | 启用 / 编辑 / 复制 / 测试连通 / 删除 全图标化(lucide-react) | -| **模型映射** | 别名("provider/model" 形式)↔ 实际请求模型 ID,自定义显示名、上下文长度、1M 上下文声明、能力 | +| **模型映射** | 别名("provider/model" 形式)↔ 实际请求模型 ID,自定义显示名、上下文长度、最大输出 Token(`max_output_size`)、能力(仅 Kimi Code) | | **自动上下文** | 拉取模型时 **API 返回 > models.dev ref > 正则兜底** 三级优先级自动适配 | | **能力自动推导** | `image_in / video_in / tool_use` 全部由 models.dev 推得;UI 仅暴露 `thinking` 一个手动开关 | | **全局设置** | `[thinking]` 表完整支持(enabled / effort / keep),仅 Kimi Code 生效 | @@ -111,7 +111,7 @@ ![编辑供应商-模型映射](docs/screenshots/provider-model-mapping.png) -一张表管理全部模型映射:显示名、实际请求模型、上下文长度、1M 上下文声明、能力(仅"思考")、设为默认、删除。 +一张表管理全部模型映射:显示名、实际请求模型、上下文长度、最大输出 Token(`max_output_size`,可点参考值填 models.dev 上限)、能力(仅"思考")、设为默认、删除(后两项仅 Kimi Code)。 **用量仪表盘** diff --git a/README_EN.md b/README_EN.md index c36fa7a..6d6904a 100644 --- a/README_EN.md +++ b/README_EN.md @@ -73,7 +73,7 @@ Build instructions: [`docs/BUILD.md`](./docs/BUILD.md). | **Connectivity test** | Real `base_url` latency, coloured bubble (green / orange / red), auto-dismiss in 6 seconds | | **Duplicate provider** | Deep-copy a provider + all its models; key auto-suffixed to `xxx-copy` | | **Iconified actions** | Activate / Edit / Duplicate / Test Connectivity / Delete via lucide-react | -| **Model mapping** | Alias (`"provider/model"`) ↔ real model ID, with display name, context size, 1M-context flag, and capabilities | +| **Model mapping** | Alias (`"provider/model"`) ↔ real model ID, with display name, context size, max output tokens (`max_output_size`), and capabilities (Kimi Code only) | | **Auto context size** | On model fetch: **API response > models.dev ref > regex fallback** — three-tier priority | | **Auto capabilities** | `image_in / video_in / tool_use` all derived from models.dev; UI only exposes `thinking` as a manual toggle | | **Global settings** | Full `[thinking]` table (enabled / effort / keep); Kimi Code only | @@ -105,7 +105,7 @@ Provider name, notes, official URL, managed-provider toggle, API format, API key ![Edit provider — model mapping](docs/screenshots/provider-model-mapping.png) -A single table for all model mappings: display name, real model ID, context length, 1M-context flag, capability (thinking only), default toggle, delete. +A single table for all model mappings: display name, real model ID, context length, max output tokens (`max_output_size`, with a click-to-fill models.dev reference), capability (thinking only), default toggle, delete (the last two are Kimi Code only). **Usage dashboard** @@ -175,7 +175,7 @@ Per-workspace Kimi Code session browsing, active / archived / all filters; strea - **`config.toml` is the authoritative source for Kimi Code**: all providers and models are always written in full; `default_model` selects the active one (matching the CLI's native `/provider` behaviour). Switching only changes `default_model`; newly added providers are auto-promoted to the top of the list and never get overwritten - **SQLite holds Kimi Switch-private metadata only**: notes, official URLs, per-agent remembered default model (the `settings` table), ordering. Theme / language / last-update-check live in frontend `localStorage` (WebView2), not under `~/.kimi-switch`. It acts as a fallback when `config.toml` is incomplete - **`raw_other` passes unknown fields through untouched**, including `[oauth]` blocks — round-trips never drop fields -- **models.dev snapshot**: derived from `https://models.dev/api.json`, cached to a local JSON; `capabilitiesFromRef` derives `thinking / image_in / video_in / tool_use`, `getModelRef` derives `max_context_size / display_name` +- **models.dev snapshot**: derived from `https://models.dev/api.json`, cached to a local JSON; `capabilitiesFromRef` derives `thinking / image_in / video_in / tool_use`, `getModelRef` derives `max_context_size / display_name / max output tokens` - **Override env vars**: `KIMI_CODE_HOME` / `PI_CODING_AGENT_DIR` override the Kimi Code / Pi dirs; Kimi Switch's own data dir is fixed at `~/.kimi-switch` (no env override yet). See [Data Storage Locations](#data-storage-locations) ## Feature Details diff --git a/docs/release-notes/release-notes-v0.8.0.md b/docs/release-notes/release-notes-v0.8.0.md new file mode 100644 index 0000000..464511b --- /dev/null +++ b/docs/release-notes/release-notes-v0.8.0.md @@ -0,0 +1,14 @@ +# KimiSwitch v0.8.0 + +## 新功能 + +- **模型映射表新增「最大输出 Token」列**(Kimi Code):为每个模型单独设置 `max_output_size`,写入 config.toml 的 `[models.""]`,留空则用上游默认值。models.dev 收录了该模型的输出上限时,输入框下方显示可点击的参考值(如 `参考 128000`),一键填入;「获取模型列表」添加模型时也会自动预填参考值 +- **models.dev 快照新增 output 上限字段**:内置快照同步至 2026-10-03(8,385 模型 / 226 供应商,97.4% 模型带 output 上限);应用内「同步 models.dev」(Rust 侧)同步支持该字段,同步后参考值不丢失 + +## 修复 + +- **移除无效的「声明支持 1M」复选框**(Kimi Code):该控件在读取时由上下文长度派生、写入时不持久化,勾选永远不会生效,已由「最大输出 Token」列取代 + +## 备注 + +- 前端 + 抓取脚本 + Rust 快照构建逻辑;新增 `max_output_size` 读写纯函数 9 个单测 + config.toml 全链路往返测试,全量测试通过(前端 126 / Rust 162) diff --git a/package.json b/package.json index 5374b41..1e5e351 100644 --- a/package.json +++ b/package.json @@ -1,7 +1,7 @@ { "name": "kimiswitch", "private": true, - "version": "0.7.23", + "version": "0.8.0", "type": "module", "scripts": { "dev": "vite", diff --git a/scripts/fetch-models-dev.mjs b/scripts/fetch-models-dev.mjs index 94f384c..46e1249 100644 --- a/scripts/fetch-models-dev.mjs +++ b/scripts/fetch-models-dev.mjs @@ -133,6 +133,11 @@ async function main() { if (typeof m.limit?.context === "number" && m.limit.context > 0) { entry.context = m.limit.context; } + // Max output tokens — the reference value behind the max_output_size + // column (model mapping table). + if (typeof m.limit?.output === "number" && m.limit.output > 0) { + entry.output = m.limit.output; + } if (m.reasoning === true) entry.reasoning = true; if (m.tool_call === true) entry.tool_call = true; if (m.structured_output === true) entry.structured_output = true; diff --git a/src-tauri/Cargo.lock b/src-tauri/Cargo.lock index b5e9d9e..91d0b2f 100644 --- a/src-tauri/Cargo.lock +++ b/src-tauri/Cargo.lock @@ -1978,7 +1978,7 @@ dependencies = [ [[package]] name = "kimiswitch" -version = "0.7.23" +version = "0.8.0" dependencies = [ "anyhow", "chrono", diff --git a/src-tauri/Cargo.toml b/src-tauri/Cargo.toml index aa14806..48c8136 100644 --- a/src-tauri/Cargo.toml +++ b/src-tauri/Cargo.toml @@ -1,6 +1,6 @@ [package] name = "kimiswitch" -version = "0.7.23" +version = "0.8.0" description = "Kimi Switch - model config manager" authors = ["codingplan.site"] edition = "2021" diff --git a/src-tauri/src/kimi_code_io.rs b/src-tauri/src/kimi_code_io.rs index 4b4f0c4..21c9d6f 100644 --- a/src-tauri/src/kimi_code_io.rs +++ b/src-tauri/src/kimi_code_io.rs @@ -648,6 +648,114 @@ api_key = "" assert!(caps.iter().any(|v| v.as_str() == Some("thinking"))); } + #[test] + fn kimi_code_export_writes_and_drops_max_output_size() { + // max_output_size is not a first-class Model field: it round-trips + // through raw_other and must land in config.toml verbatim, and must be + // removable (unset = upstream default) without disturbing its siblings. + let mut providers = IndexMap::new(); + providers.insert( + "p".to_string(), + Provider { + name: "p".to_string(), + provider_type: ProviderType::Openai, + base_url: Some("https://a.example.com".to_string()), + api_key: Some("sk-a".to_string()), + api_key_env: None, + env: IndexMap::new(), + note: None, + official_url: None, + managed: false, + enabled: true, + active: true, + icon: None, + icon_color: None, + raw_other: Value::Null, + usage_kinds: None, + usage_config: None, + }, + ); + let make_model = |raw_other: Value| Model { + alias: "p/m".to_string(), + provider: "p".to_string(), + model: "m".to_string(), + max_context_size: 128_000, + display_name: None, + supports_1m: false, + capabilities: vec![], + raw_other, + }; + + let mut models = IndexMap::new(); + models.insert( + "p/m".to_string(), + make_model(serde_json::json!({ + "max_output_size": 32768, + "support_efforts": ["low", "high"], + })), + ); + let config = Config { + default_model: None, + providers, + models, + raw_other: Value::Null, + imported_section_keys: Vec::new(), + }; + + let exported = config_to_kimi_code(&config, None); + let root = exported.as_table().unwrap(); + let model = root + .get("models") + .unwrap() + .as_table() + .unwrap() + .get("p/m") + .unwrap() + .as_table() + .unwrap(); + assert_eq!(model.get("max_output_size").and_then(|v| v.as_integer()), Some(32768)); + assert!(model.get("support_efforts").is_some()); + + // Re-import through real TOML text (the same serialize → parse path + // save_kimi_code_config / load_kimi_code_config use): the key must + // come back in the model's raw_other as an integer. + let toml_text = toml::to_string_pretty(&exported).unwrap(); + assert!(toml_text.contains("max_output_size = 32768"), "toml: {toml_text}"); + let reparsed: TomlValue = toml_text.parse().unwrap(); + let parsed = kimi_code_to_config(&reparsed); + let loaded = parsed.models.get("p/m").unwrap(); + assert_eq!( + loaded.raw_other.get("max_output_size").and_then(|v| v.as_i64()), + Some(32768) + ); + + // Clearing the override drops the key entirely. + let mut models = IndexMap::new(); + models.insert( + "p/m".to_string(), + make_model(serde_json::json!({ "support_efforts": ["low", "high"] })), + ); + let config = Config { + default_model: None, + providers: IndexMap::new(), + models, + raw_other: Value::Null, + imported_section_keys: Vec::new(), + }; + let exported = config_to_kimi_code(&config, None); + let root = exported.as_table().unwrap(); + let model = root + .get("models") + .unwrap() + .as_table() + .unwrap() + .get("p/m") + .unwrap() + .as_table() + .unwrap(); + assert!(model.get("max_output_size").is_none()); + } + #[test] fn kimi_code_export_adds_opencode_go_session_header() { let make_provider = |base_url: &str, raw_other: Value| Provider { diff --git a/src-tauri/src/models_dev.rs b/src-tauri/src/models_dev.rs index fea9091..288440c 100644 --- a/src-tauri/src/models_dev.rs +++ b/src-tauri/src/models_dev.rs @@ -138,6 +138,14 @@ fn build_snapshot(raw: &Value) -> Result { { entry.insert("context".into(), num(ctx)); } + // Max output tokens — reference for the max_output_size column. + if let Some(out) = m + .pointer("/limit/output") + .and_then(|x| x.as_f64()) + .filter(|o| *o > 0.0) + { + entry.insert("output".into(), num(out)); + } for flag in ["reasoning", "tool_call", "structured_output"] { if m.get(flag).and_then(|x| x.as_bool()) == Some(true) { entry.insert(flag.into(), Value::Bool(true)); @@ -286,7 +294,7 @@ mod tests { }, "kimi-image": { "name": "Kimi Image", - "limit": { "context": 0 }, + "limit": { "context": 0, "output": 0 }, "modalities": { "input": ["text", "image", "video"] } } } @@ -309,6 +317,7 @@ mod tests { let k3 = &obj["moonshotai/kimi-k3"]; assert_eq!(k3["name"], "Kimi K3"); assert_eq!(k3["context"], 262144); + assert_eq!(k3["output"], 16384); assert_eq!(k3["reasoning"], true); assert_eq!(k3["tool_call"], true); assert_eq!(k3["structured_output"], true); @@ -318,9 +327,10 @@ mod tests { assert_eq!(k3["cost"]["cache_read"], 0.3); assert_eq!(k3["cost"]["cache_write"], 0.0); - // Context 0 → dropped; video modality → flagged. + // Context 0 / output 0 → dropped; video modality → flagged. let img = &obj["moonshotai/kimi-image"]; assert!(img.get("context").is_none()); + assert!(img.get("output").is_none()); assert_eq!(img["video"], true); assert!(img.get("cost").is_none()); diff --git a/src-tauri/tauri.conf.json b/src-tauri/tauri.conf.json index d7d93ce..8a15594 100644 --- a/src-tauri/tauri.conf.json +++ b/src-tauri/tauri.conf.json @@ -1,6 +1,6 @@ { "productName": "Kimi Switch", - "version": "0.7.23", + "version": "0.8.0", "identifier": "com.kimiswitch.app", "build": { "beforeDevCommand": "npm run dev", diff --git a/src/components/ProviderEdit.tsx b/src/components/ProviderEdit.tsx index 9e8f35a..229d345 100644 --- a/src/components/ProviderEdit.tsx +++ b/src/components/ProviderEdit.tsx @@ -1,11 +1,17 @@ -import { useEffect, useId, useMemo, useState } from "react"; +import { useEffect, useId, useMemo, useReducer, useRef, useState } from "react"; import { createPortal } from "react-dom"; import { invoke } from "@tauri-apps/api/core"; import { useTranslation } from "../i18n"; import type { TranslationKey } from "../i18n/zh"; import { findPresetForProvider } from "../config/providerPresets"; import { getDefaultMaxContextSize } from "../lib/model-defaults"; -import { capabilitiesFromRef, getModelRef } from "../lib/models-dev"; +import { + parseMaxOutputInput, + readMaxOutputSize, + setMaxOutputSize, + withMaxOutputSize, +} from "../lib/model-max-output"; +import { capabilitiesFromRef, getModelRef, modelsDevReady } from "../lib/models-dev"; import { getIconMetadata } from "../icons/extracted/metadata"; import { AgentSettingsPanel } from "./AgentSettingsPanel"; import { KimiOAuthDialog } from "./KimiOAuthDialog"; @@ -567,6 +573,33 @@ function ModelMapping({ const [discoverError, setDiscoverError] = useState(null); const [fetchThinking, setFetchThinking] = useState(true); + // The models.dev index loads in the background; re-render once it lands so + // the max-output "参考" hints appear even when this panel mounted first. + const [, forceModelsDevReady] = useReducer((x: number) => x + 1, 0); + // One-shot backfill: models with no max_output_size yet get the models.dev + // `output` cap filled in automatically. onModelChange only mutates the + // in-memory config — the user still reviews and presses 保存配置. + const backfillDone = useRef(false); + const modelsRef = useRef(models); + modelsRef.current = models; + useEffect(() => { + let alive = true; + modelsDevReady().then(() => { + if (!alive) return; + forceModelsDevReady(); + if (backfillDone.current || agent !== "kimi_code") return; + backfillDone.current = true; + for (const m of modelsRef.current) { + if (readMaxOutputSize(m.raw_other) !== undefined) continue; + const output = getModelRef(m.model)?.output; + if (output !== undefined) onModelChange(withMaxOutputSize(m, output)); + } + }); + return () => { + alive = false; + }; + }, []); + const handleDiscover = async () => { setDiscovering(true); setDiscoverError(null); @@ -636,6 +669,12 @@ function ModelMapping({ : fetchThinking ? ["thinking"] : [], + // Seed the max_output_size override only when models.dev knows the + // cap — an absent key means the upstream default applies. kimi_code + // only: Pi's equivalent is `maxTokens` (not written by this UI). + ...(agent === "kimi_code" && ref?.output + ? { raw_other: setMaxOutputSize(undefined, ref.output) } + : {}), }); } onBulkAdd(toAdd); @@ -648,7 +687,10 @@ function ModelMapping({

{t("modelMapping")}

-

{t("modelMappingDesc")}

+

+ {t("modelMappingDesc")} + {agent === "kimi_code" ? ` ${t("maxOutputSizeDesc")}` : ""} +

+ )} +
+ ); +} diff --git a/src/components/SubagentSettingsPage.tsx b/src/components/SubagentSettingsPage.tsx index 4daaba3..c96590d 100644 --- a/src/components/SubagentSettingsPage.tsx +++ b/src/components/SubagentSettingsPage.tsx @@ -45,7 +45,7 @@ interface SubagentSettingsPageProps { // Effort tiers the CLI accepts; the value is forwarded verbatim upstream, and // per-model support comes from `[models.""] support_efforts`. -const EFFORTS = ["low", "medium", "high", "max", "xhigh"] as const; +const EFFORTS = ["low", "medium", "high", "xhigh", "max"] as const; const EFFORT_LABELS: Record<(typeof EFFORTS)[number], TranslationKey> = { low: "thinkingLow", medium: "thinkingMedium", diff --git a/src/i18n/en.ts b/src/i18n/en.ts index f2b7c02..2807385 100644 --- a/src/i18n/en.ts +++ b/src/i18n/en.ts @@ -111,7 +111,8 @@ export const enTranslations: Record = { baseUrlHint: "Enter a {type} API-compatible endpoint without a trailing slash.", envPairs: "Env pairs", addEnv: "+ Add", - modelMappingDesc: "Display name only affects the /model menu; 1M declares context capability for Kimi Code.", + modelMappingDesc: "Display name only affects the /model menu.", + maxOutputSizeDesc: "Max output tokens writes max_output_size — leave blank to use the upstream default.", oneClickSetup: "One-click setup", fetchModels: "Fetch models", fetchingModels: "Fetching...", @@ -121,7 +122,8 @@ export const enTranslations: Record = { displayName: "Display name", actualModel: "Actual model", contextSize: "Context size", - supports1M: "Supports 1M", + maxOutputSize: "Max output tokens", + maxOutputRef: "Ref {value}", default: "Default", operation: "Operation", isDefault: "Default", diff --git a/src/i18n/zh.ts b/src/i18n/zh.ts index f49ac25..a33674b 100644 --- a/src/i18n/zh.ts +++ b/src/i18n/zh.ts @@ -108,7 +108,8 @@ export const zhTranslations = { baseUrlHint: "填写兼容 {type} API 的服务端点地址,不要以斜杠结尾", envPairs: "Env 键值对", addEnv: "+ 添加", - modelMappingDesc: "显示名称只影响 /model 菜单;1M 只是给 Kimi Code 的上下文能力声明。", + modelMappingDesc: "显示名称只影响 /model 菜单。", + maxOutputSizeDesc: "「最大输出 Token」写入 max_output_size,留空则用上游默认值。", oneClickSetup: "一键设置", fetchModels: "获取模型列表", fetchingModels: "获取中...", @@ -118,7 +119,8 @@ export const zhTranslations = { displayName: "显示名称", actualModel: "实际请求模型", contextSize: "上下文长度", - supports1M: "声明支持 1M", + maxOutputSize: "最大输出 Token", + maxOutputRef: "参考 {value}", default: "默认", operation: "操作", isDefault: "已默认", @@ -158,9 +160,9 @@ export const zhTranslations = { enableThinking: "启用思考", thinkingLevel: "思考等级", thinkingKeep: "保留思考内容", - thinkingLow: "低", - thinkingMedium: "中", - thinkingHigh: "高", + thinkingLow: "Low", + thinkingMedium: "Medium", + thinkingHigh: "High", thinkingMax: "Max", thinkingXHigh: "XHigh", thinkingEffortUnsupported: "当前默认模型不支持思考。", diff --git a/src/lib/model-max-output.test.ts b/src/lib/model-max-output.test.ts new file mode 100644 index 0000000..0413ac8 --- /dev/null +++ b/src/lib/model-max-output.test.ts @@ -0,0 +1,83 @@ +import { describe, expect, it } from "vitest"; +import { + parseMaxOutputInput, + readMaxOutputSize, + setMaxOutputSize, + withMaxOutputSize, +} from "./model-max-output"; + +const model = (raw_other: unknown) => + ({ + alias: "a", + provider: "p", + model: "m", + max_context_size: 0, + display_name: null, + raw_other, + }) as const; + +describe("readMaxOutputSize", () => { + it("reads the key from a model entry's preserved config fields", () => { + expect(readMaxOutputSize({ max_output_size: 32768 })).toBe(32768); + }); + + it("treats absent / malformed / non-positive values as unset", () => { + expect(readMaxOutputSize(undefined)).toBeUndefined(); + expect(readMaxOutputSize(null)).toBeUndefined(); + expect(readMaxOutputSize({})).toBeUndefined(); + expect(readMaxOutputSize("32768")).toBeUndefined(); + expect(readMaxOutputSize({ max_output_size: 0 })).toBeUndefined(); + expect(readMaxOutputSize({ max_output_size: -1 })).toBeUndefined(); + expect(readMaxOutputSize({ max_output_size: NaN })).toBeUndefined(); + expect(readMaxOutputSize([1, 2])).toBeUndefined(); + }); +}); + +describe("parseMaxOutputInput", () => { + it("keeps positive integers", () => { + expect(parseMaxOutputInput("32768")).toBe(32768); + }); + + it("maps blank / zero / garbage to unset", () => { + expect(parseMaxOutputInput("")).toBeUndefined(); + expect(parseMaxOutputInput("0")).toBeUndefined(); + expect(parseMaxOutputInput("-5")).toBeUndefined(); + expect(parseMaxOutputInput("abc")).toBeUndefined(); + }); +}); + +describe("setMaxOutputSize", () => { + it("sets the key while preserving other preserved fields", () => { + expect(setMaxOutputSize({ support_efforts: ["low"] }, 16384)).toEqual({ + support_efforts: ["low"], + max_output_size: 16384, + }); + }); + + it("drops the key when the value is unset", () => { + expect(setMaxOutputSize({ max_output_size: 16384, force: true }, undefined)).toEqual({ + force: true, + }); + }); + + it("starts from an empty object for null / non-object raw_other", () => { + expect(setMaxOutputSize(null, 1024)).toEqual({ max_output_size: 1024 }); + expect(setMaxOutputSize("junk", 1024)).toEqual({ max_output_size: 1024 }); + expect(setMaxOutputSize({ max_output_size: 1 }, undefined)).toEqual({}); + }); +}); + +describe("withMaxOutputSize", () => { + it("returns a new model with the override applied", () => { + const before = model({ max_output_size: 4096 }); + const after = withMaxOutputSize(before, 8192); + expect(readMaxOutputSize(after.raw_other)).toBe(8192); + expect(readMaxOutputSize(before.raw_other)).toBe(4096); + expect(after).not.toBe(before); + }); + + it("clears the override without touching other fields", () => { + const after = withMaxOutputSize(model({ max_output_size: 4096 }), undefined); + expect(after.raw_other).toEqual({}); + }); +}); diff --git a/src/lib/model-max-output.ts b/src/lib/model-max-output.ts new file mode 100644 index 0000000..018602b --- /dev/null +++ b/src/lib/model-max-output.ts @@ -0,0 +1,59 @@ +import type { Model } from "../types"; + +/** + * Helpers for the `max_output_size` model field. + * + * kimi-code keeps it in `[models.""] max_output_size`; it is not a + * first-class field of Kimi Switch's Rust `Model` struct, so it round-trips + * through `raw_other` (the pass-through bucket for unknown keys). Absent key = + * unset = the upstream default applies, so clearing the field removes the key + * rather than writing 0. + */ + +function asRecord(value: unknown): Record { + if (value && typeof value === "object" && !Array.isArray(value)) { + return value as Record; + } + return {}; +} + +/** Read the override; undefined when absent or malformed (treated as unset). */ +export function readMaxOutputSize(rawOther: unknown): number | undefined { + const value = asRecord(rawOther).max_output_size; + if (typeof value !== "number" || !Number.isFinite(value) || value <= 0) { + return undefined; + } + return value; +} + +/** Parse the number input: blank / 0 / NaN mean "unset". */ +export function parseMaxOutputInput(text: string): number | undefined { + const value = parseInt(text, 10); + if (isNaN(value) || value <= 0) return undefined; + return value; +} + +/** + * Set (positive value) or drop (undefined) `max_output_size` on a raw_other + * blob; every other preserved key is kept. + */ +export function setMaxOutputSize( + rawOther: unknown, + value: number | undefined, +): Record { + const next = { ...asRecord(rawOther) }; + if (value === undefined) { + delete next.max_output_size; + } else { + next.max_output_size = Math.floor(value); + } + return next; +} + +/** Immutable update of a model entry's max_output_size override. */ +export function withMaxOutputSize( + model: Model, + value: number | undefined, +): Model { + return { ...model, raw_other: setMaxOutputSize(model.raw_other, value) }; +} diff --git a/src/lib/models-dev.ts b/src/lib/models-dev.ts index 19ebf66..f65f388 100644 --- a/src/lib/models-dev.ts +++ b/src/lib/models-dev.ts @@ -25,6 +25,8 @@ export interface ModelCost { export interface ModelRef { name?: string; context?: number; + /** Upper bound on generated tokens (`limit.output` on models.dev). */ + output?: number; reasoning?: boolean; tool_call?: boolean; structured_output?: boolean; diff --git a/src/lib/thinking-efforts.ts b/src/lib/thinking-efforts.ts index 1778b21..8c15e4b 100644 --- a/src/lib/thinking-efforts.ts +++ b/src/lib/thinking-efforts.ts @@ -2,7 +2,7 @@ * Thinking-effort capability resolution. * * kimi-code's `[thinking] effort` is a free-form string (low/medium/high/ - * max/xhigh in practice) that the CLI forwards verbatim to OpenAI-compatible + * xhigh/max in practice) that the CLI forwards verbatim to OpenAI-compatible * upstreams as `reasoning_effort`. Which tiers a given model accepts is * declared per model in config.toml (`[models.""] support_efforts`); * models.dev only carries a boolean `reasoning` flag with no tier list. The @@ -16,8 +16,8 @@ export const THINKING_EFFORTS = [ "low", "medium", "high", - "max", "xhigh", + "max", ] as const; export type ThinkingEffort = (typeof THINKING_EFFORTS)[number];