diff --git a/README.en.md b/README.en.md index 69c55cd..5334878 100644 --- a/README.en.md +++ b/README.en.md @@ -111,6 +111,31 @@ The MCP server starts automatically with the session. | Multimodal model | Paste with Alt+V directly — native vision, plugin not involved | | Non-multimodal + image path | Type `@screenshot.png` or paste the path; the model calls `read_image` | | Non-multimodal + just screenshotted/copied | Do **not** Alt+V paste (the CLI rejects pasting on non-multimodal models with `Current model does not support image input`); just ask "analyze this screenshot" — the model calls `read_clipboard_image` to read the system clipboard | +| Non-multimodal + paste was rejected | Tell the model "my paste was blocked" — it will switch to `read_clipboard_image` automatically; no need to save the file | + +#### Want Alt+V pasting? (declare `image_in` on the model) + +Kimi Code's frontend blocks pasting on models that lack image support. Add +`image_in` to the model in `~/.kimi-code/config.toml`: + +```toml +[models."opencode-go/deepseek-v4-flash"] +capabilities = [ "thinking", "tool_use", "image_in" ] # append image_in +``` + +> ⚠️ **Important**: `image_in` only lets the **frontend accept** the paste — it does +> not mean the provider can actually receive images. +> +> - Declare it **only when the provider truly supports image input** (e.g. +> MiniMax-M3, k3, gpt-5.6-luna, grok-4.5) — those models see natively and the +> plugin stays idle +> - **Never declare it on text-only providers** (e.g. deepseek-v4-flash, GLM text +> variants): the paste gets accepted, then the request fails at the provider with +> `400 unknown variant image_url, expected text`. The frontend block is a guard. +> +> Correct usage for text-only models: `@image-path` (`read_image`) or screenshot and +> ask directly (`read_clipboard_image`); if a paste is rejected, tell the model +> "the paste was blocked" and it will read the clipboard instead. **No commands needed.** ## Activation and triggering diff --git a/README.md b/README.md index 6ef5636..fd3051f 100644 --- a/README.md +++ b/README.md @@ -95,7 +95,12 @@ KimiCode 前端默认会拦截「不支持图片输入」模型的粘贴。在 ` capabilities = [ "thinking", "tool_use", "image_in" ] # 追加 image_in ``` -> 注意:声明 `image_in` 只是让前端放行粘贴。纯文本模型依然**看不懂**图片内容——这正是插件的用武之地:模型收到图片但看不见内容时,会按 SYSTEM.md 规则自动调 `read_clipboard_image` 从剪贴板读图。**全程零命令、零前缀**。 +> ⚠️ **重要警告**:`image_in` 只是让**前端放行**,不代表 provider 真能收图。 +> +> - **仅当 provider 本身支持图片输入**(如 MiniMax-M3、k3、gpt-5.6-luna、grok-4.5 等)时声明才有意义——这些模型粘贴后原生看图,插件闲置 +> - **纯文本 provider**(如 deepseek-v4-flash、glm 文本版)**不要声明**——声明后粘贴会被放行,但请求发给 provider 时图片 part 会被拒(400 `unknown variant image_url, expected text`)。此时前端拦截反而是保护 +> +> 纯文本模型的正确用法:`@图片路径`(read_image)或截图后直接提问(read_clipboard_image 读剪贴板);粘贴被拦截就直接告诉模型「粘贴被拦截了」,它会自动改读剪贴板。**全程零命令**。 ## 激活与触发 diff --git a/package.json b/package.json index 58f218a..e14c9f2 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "kimi-eyes", - "version": "1.0.4", + "version": "1.0.5", "description": "Give non-multimodal models in Kimi Code the ability to analyze images and screenshots via a user-configured VLM (OpenAI-compatible or Anthropic protocol).", "keywords": [ "kimi",