diff --git a/README.en.md b/README.en.md index 98b5e5a..930612c 100644 --- a/README.en.md +++ b/README.en.md @@ -113,6 +113,17 @@ The MCP server starts automatically with the session. | Non-multimodal + just screenshotted/copied | Do **not** Alt+V paste (the CLI rejects pasting on non-multimodal models with `Current model does not support image input`); just ask "analyze this screenshot" — the model calls `read_clipboard_image` to read the system clipboard | | Non-multimodal + paste was rejected | Tell the model "my paste was blocked" — it will switch to `read_clipboard_image` automatically; no need to save the file | +#### Ways to view images × `image_in` + +| Way | Needs `image_in`? | Text-only OK? | Effort | +| --- | --- | --- | --- | +| Ask right after a screenshot (auto clipboard) | No | ✅ | Easiest | +| `@path` / give a path | No | ✅ | Easy | +| `/skill kimi-eyes` + Alt-V | Yes | ✅ | More steps | +| Plain Alt-V | Declare or not | ❌ | Won't work | + +> For text-only models, prefer the first two day-to-day; only reach for `/skill kimi-eyes` when you really want the paste gesture (see below). + #### Want Alt+V pasting? (declare `image_in` on the model) Kimi Code's frontend blocks pasting on models that lack image support. Add @@ -137,6 +148,37 @@ capabilities = [ "thinking", "tool_use", "image_in" ] # append image_in > ask directly (`read_clipboard_image`); if a paste is rejected, tell the model > "the paste was blocked" and it will read the clipboard instead. **No commands needed.** +#### Text-only model but you really want to paste? Use `/skill kimi-eyes` + +The warning above says: on a text-only provider, declaring `image_in` and using **Alt-V paste** sends an `image_url` part to the provider and triggers a 400. But the same `image_in` declaration is harmless if you go through the **`/skill` command** instead — `/skill` renders the pasted image as an `Attached image file: ` **plain-text path**, producing no image part. This plugin ships a skill that exploits exactly this channel. + +**One-time setup (two steps)**: + +1. Declare `image_in` on the text-only model (only to pass the `/skill` frontend check; `/skill` sends no image part, so **no 400**): + +```toml +[models."opencode-go/deepseek-v4-flash"] +capabilities = [ "thinking", "tool_use", "image_in" ] # append image_in +``` + +2. Register the skill — add the plugin's `skills/` dir to the scan list in `~/.kimi-code/config.toml`: + +```toml +extra_skill_dirs = [ "D:/AIGC/Plugin/kimi-eyes/skills" ] +``` + +> If the plugin lives elsewhere, use its actual `kimi-eyes/skills` path. After restarting the session, `/skill kimi-eyes` appears in the `/` completion menu. + +**Usage**: + +``` +/skill kimi-eyes what does this chart show? ← then Alt-V paste the image, press Enter +``` + +The model extracts the path from `Attached image file: `, calls kimi-eyes' `read_image` (the external VLM returns a text description), then answers. No image part is ever produced, so text-only models never 400. + +> ⚠️ **Constraint after declaring `image_in`**: for image tasks always go through `/skill kimi-eyes`; do **not** Alt-V paste directly (the main-prompt channel still sends an image part to text-only providers → 400). `@image-path` and "ask after screenshot" keep working as before. + ## Activation and triggering ### Activation (one-time) diff --git a/README.md b/README.md index 34e399d..4b50bb2 100644 --- a/README.md +++ b/README.md @@ -83,9 +83,20 @@ MCP 服务器会随会话自动启动。 | 多模态模型 | 直接 Alt+V 粘贴图片,原生看图,插件不参与 | | 非多模态 + 有图片路径 | 输入 `@截图.png` 或直接给路径,模型自动调 `read_image` | | 非多模态 + 刚截图/复制 | 直接提问「分析这张截图」,模型自动调 `read_clipboard_image` 读系统剪贴板 | -| 非多模态 + 想用 Alt+V 粘贴 | 需先在 config.toml 给模型声明 `image_in`(见下),粘贴后直接提问即可——模型看不懂图片内容时会自动调 `read_clipboard_image`(剪贴板仍保留原图) | +| 非多模态 + 想用 Alt+V 粘贴 | 纯文本模型需 `/skill kimi-eyes` 配合(见下),单纯粘贴会报错 | | 粘贴被拦截 | 直接告诉模型「我粘贴图片被拦截了」,模型会自动改读剪贴板,无需你保存文件 | +#### 看图方式 × `image_in` 速查 + +| 看图方式 | 需要 `image_in`? | 纯文本可用? | 顺滑度 | +| --- | --- | --- | --- | +| 截图后随口一问(自动读剪贴板) | 否 | ✅ | 最省事 | +| `@路径` / 给路径 | 否 | ✅ | 省事 | +| `/skill kimi-eyes` + Alt-V | 是 | ✅ | 多步 | +| 单纯 Alt-V | 声明与否 | ❌ | 走不通 | + +> 纯文本模型日常看图优先用前两种;只有「就想用粘贴这个动作」时才走 `/skill kimi-eyes`(见下)。 + #### 想用 Alt+V 粘贴?(给模型声明 `image_in`) KimiCode 前端默认会拦截「不支持图片输入」模型的粘贴。在 `~/.kimi-code/config.toml` 给对应模型加上 `image_in` 即可放行: @@ -102,6 +113,37 @@ capabilities = [ "thinking", "tool_use", "image_in" ] # 追加 image_in > > 纯文本模型的正确用法:`@图片路径`(read_image)或截图后直接提问(read_clipboard_image 读剪贴板);粘贴被拦截就直接告诉模型「粘贴被拦截了」,它会自动改读剪贴板。**全程零命令**。 +#### 纯文本模型也想「粘贴看图」?用 `/skill`(推荐) + +上面警告过:纯文本 provider 声明 `image_in` 后走 **Alt-V 粘贴**,图片会作为 `image_url` part 发给 provider 触发 400。但同样声明 `image_in`,改走 **`/skill` 命令**就不会 400——`/skill` 通道把图片渲染成 `Attached image file: <路径>` 的**纯文本路径**,不产生 image part。本插件自带一个 skill 利用这条通道。 + +**一次性配置(两步)**: + +1. 给纯文本模型声明 `image_in`(仅为通过 `/skill` 的前端校验;`/skill` 下图片不走 part,**不会**触发 400): + +```toml +[models."opencode-go/deepseek-v4-flash"] +capabilities = [ "thinking", "tool_use", "image_in" ] # 追加 image_in +``` + +2. 注册 skill——在 `~/.kimi-code/config.toml` 把插件目录的 `skills/` 加入扫描: + +```toml +extra_skill_dirs = [ "D:/AIGC/Plugin/kimi-eyes/skills" ] +``` + +> 插件装在别处就换成实际的 `kimi-eyes/skills` 目录。重启会话后 `/skill kimi-eyes` 会出现在 `/` 补全菜单。 + +**用法**: + +``` +/skill kimi-eyes 这张图表说明了什么? ← 然后 Alt-V 粘贴图片,回车 +``` + +模型会从消息里的 `Attached image file: <路径>` 提取路径,调用 kimi-eyes 的 `read_image`(外部 VLM 返回文字描述)再作答。全程不产生 image part,纯文本模型不会 400。 + +> ⚠️ **声明 `image_in` 后的使用约束**:看图一律走 `/skill kimi-eyes`,**不要**再直接 Alt-V 粘贴(主 prompt 通道仍会把图片作 image part 发给纯文本 provider → 400)。`@图片路径` 和「截图后提问」两种老用法不受影响。 + ## 激活与触发 ### 激活(一次性) diff --git a/docs/index.html b/docs/index.html index 229943f..31ee88f 100644 --- a/docs/index.html +++ b/docs/index.html @@ -429,7 +429,10 @@ npx kimi-eyes setup # 使用 @截图.png 这个图里有什么? # → read_image -分析这张截图 # → read_clipboard_image +分析这张截图 # → read_clipboard_image + +# 纯文本模型想「粘贴」看图(需先声明 image_in + 注册 skill,详见 README): +/skill kimi-eyes 这张图说明什么? # 然后 Alt-V 粘贴图片,回车 @@ -456,6 +459,33 @@ claude mcp remove kimi-eyes + +
+
+

看图方式 × image_in

+

纯文本主模型下,不同看图姿势对 image_in 的要求和可用性不同。

+
+ + + + + + + + + + + + + + + +
看图方式需要 image_in?纯文本可用?顺滑度
截图后随口一问(自动读剪贴板)否✅最省事
@路径 / 给路径否✅省事
/skill kimi-eyes + Alt-V是✅多步
单纯 Alt-V声明与否❌走不通
+
+

日常优先用前两种;只有「就想粘贴」时才走 /skill kimi-eyes。

+
+
+
diff --git a/kimi.plugin.json b/kimi.plugin.json index b039afa..69cc3ff 100644 --- a/kimi.plugin.json +++ b/kimi.plugin.json @@ -1,6 +1,6 @@ { "name": "kimi-eyes", - "version": "1.0.6", + "version": "1.0.7", "description": "Give non-multimodal models vision: analyze local images and clipboard screenshots through a user-configured VLM API (OpenAI-compatible or Anthropic protocol).", "keywords": [ "vision", diff --git a/package.json b/package.json index 559e911..1752bdf 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "kimi-eyes", - "version": "1.0.6", + "version": "1.0.7", "description": "Give non-multimodal models in Kimi Code the ability to analyze images and screenshots via a user-configured VLM (OpenAI-compatible or Anthropic protocol).", "keywords": [ "kimi", diff --git a/skills/kimi-eyes/SKILL.md b/skills/kimi-eyes/SKILL.md new file mode 100644 index 0000000..7d817c2 --- /dev/null +++ b/skills/kimi-eyes/SKILL.md @@ -0,0 +1,36 @@ +--- +name: kimi-eyes +description: 纯文本主模型的看图入口。通过 /skill kimi-eyes 触发,把粘贴的图片交给 kimi-eyes 配置的外部 VLM 分析,返回文字描述。 +type: prompt +whenToUse: 当用户用 /skill kimi-eyes 发送图片、且当前主模型不具备原生看图能力时 +--- + +# Kimi Eyes — 纯文本主模型看图 + +## 为什么需要这个 skill + +`/skill` 通道下,用户粘贴的图片以 `Attached image file: <缓存路径>` 的**纯文本**形式 +进入对话,不会作为 `image_url` content part 发给主模型。因此**纯文本主模型也不会触发 +`400 unknown variant image_url` 错误**。本 skill 利用这条文本通道,把图片交给 kimi-eyes +已配置的外部 VLM,拿回文字描述后作答。 + +## 触发条件 + +- **仅当**用户当前消息以 `/skill kimi-eyes` 开头,并附带图片 +- 用户直接 Alt-V 粘贴图片(没有 `/skill` 前缀)时**不要**走本 skill,按 kimi-eyes 常规机制处理(@图片路径 或 截图后直接提问) +- 多模态主模型(已声明 image_in 且 provider 真支持图片)原生看图即可,不需要本 skill + +## 执行步骤 + +1. 在用户消息里定位 `Attached image file: <路径>` 文本,提取其中的图片**绝对路径**(缓存目录下的文件) +2. 调用 kimi-eyes 的 `read_image` 工具: + - `path` = 第 1 步提取到的图片路径 + - `prompt` = 用户的问题;若用户没给具体问题,使用「详细描述这张图片的内容」 +3. **不要**使用 `ReadMediaFile`(它对纯文本模型无意义,且会把图片作为 part 回灌,可能再次触发 400) +4. `read_image` 返回的是外部 VLM 的**文字描述**,基于这段描述回答用户的问题 + +## 失败兜底 + +- 消息里没有 `Attached image file:` 路径 → 提示用户:「请输入 `/skill kimi-eyes` 后再粘贴图片」 +- `read_image` 报「Vision API is not configured」→ 提示用户运行 `npx kimi-eyes setup` 配置视觉 API +- `read_image` 报文件不存在 → 提示用户重新粘贴图片(缓存文件可能已失效)