feat: 新增 /skill kimi-eyes 触发入口,规避纯文本模型 image_url 400
利用 kimi-code 的 /skill 文本通道(rewriteMediaPlaceholders 把图片渲染成 Attached image file 路径,不产生 image part),让纯文本主模型也能粘贴看图:用户 /skill kimi-eyes + Alt-V 粘贴,模型从路径调 read_image 拿回文字描述。新增 skills/kimi-eyes/SKILL.md;README(中英)+ 网页新增「看图方式 × image_in」对比表与 /skill 配置说明。1.0.6 → 1.0.7。
This commit is contained in:
1 parent
64a9807fb9
commit
43f09c83f4
6 files changed
+154
-4
No files matched your search
@@ -113,6 +113,17 @@ The MCP server starts automatically with the session.
|
||||
| Non-multimodal + just screenshotted/copied | Do **not** Alt+V paste (the CLI rejects pasting on non-multimodal models with `Current model does not support image input`); just ask "analyze this screenshot" — the model calls `read_clipboard_image` to read the system clipboard |
|
||||
| Non-multimodal + paste was rejected | Tell the model "my paste was blocked" — it will switch to `read_clipboard_image` automatically; no need to save the file |
|
||||
|
||||
#### Ways to view images × `image_in`
|
||||
|
||||
| Way | Needs `image_in`? | Text-only OK? | Effort |
|
||||
| --- | --- | --- | --- |
|
||||
| Ask right after a screenshot (auto clipboard) | No | ✅ | Easiest |
|
||||
| `@path` / give a path | No | ✅ | Easy |
|
||||
| `/skill kimi-eyes` + Alt-V | Yes | ✅ | More steps |
|
||||
| Plain Alt-V | Declare or not | ❌ | Won't work |
|
||||
|
||||
> For text-only models, prefer the first two day-to-day; only reach for `/skill kimi-eyes` when you really want the paste gesture (see below).
|
||||
|
||||
#### Want Alt+V pasting? (declare `image_in` on the model)
|
||||
|
||||
Kimi Code's frontend blocks pasting on models that lack image support. Add
|
||||
@@ -137,6 +148,37 @@ capabilities = [ "thinking", "tool_use", "image_in" ] # append image_in
|
||||
> ask directly (`read_clipboard_image`); if a paste is rejected, tell the model
|
||||
> "the paste was blocked" and it will read the clipboard instead. **No commands needed.**
|
||||
|
||||
#### Text-only model but you really want to paste? Use `/skill kimi-eyes`
|
||||
|
||||
The warning above says: on a text-only provider, declaring `image_in` and using **Alt-V paste** sends an `image_url` part to the provider and triggers a 400. But the same `image_in` declaration is harmless if you go through the **`/skill` command** instead — `/skill` renders the pasted image as an `Attached image file: <path>` **plain-text path**, producing no image part. This plugin ships a skill that exploits exactly this channel.
|
||||
|
||||
**One-time setup (two steps)**:
|
||||
|
||||
1. Declare `image_in` on the text-only model (only to pass the `/skill` frontend check; `/skill` sends no image part, so **no 400**):
|
||||
|
||||
```toml
|
||||
[models."opencode-go/deepseek-v4-flash"]
|
||||
capabilities = [ "thinking", "tool_use", "image_in" ] # append image_in
|
||||
```
|
||||
|
||||
2. Register the skill — add the plugin's `skills/` dir to the scan list in `~/.kimi-code/config.toml`:
|
||||
|
||||
```toml
|
||||
extra_skill_dirs = [ "D:/AIGC/Plugin/kimi-eyes/skills" ]
|
||||
```
|
||||
|
||||
> If the plugin lives elsewhere, use its actual `kimi-eyes/skills` path. After restarting the session, `/skill kimi-eyes` appears in the `/` completion menu.
|
||||
|
||||
**Usage**:
|
||||
|
||||
```
|
||||
/skill kimi-eyes what does this chart show? ← then Alt-V paste the image, press Enter
|
||||
```
|
||||
|
||||
The model extracts the path from `Attached image file: <path>`, calls kimi-eyes' `read_image` (the external VLM returns a text description), then answers. No image part is ever produced, so text-only models never 400.
|
||||
|
||||
> ⚠️ **Constraint after declaring `image_in`**: for image tasks always go through `/skill kimi-eyes`; do **not** Alt-V paste directly (the main-prompt channel still sends an image part to text-only providers → 400). `@image-path` and "ask after screenshot" keep working as before.
|
||||
|
||||
## Activation and triggering
|
||||
|
||||
### Activation (one-time)
|
||||
|
||||
@@ -83,9 +83,20 @@ MCP 服务器会随会话自动启动。
|
||||
| 多模态模型 | 直接 Alt+V 粘贴图片,原生看图,插件不参与 |
|
||||
| 非多模态 + 有图片路径 | 输入 `@截图.png` 或直接给路径,模型自动调 `read_image` |
|
||||
| 非多模态 + 刚截图/复制 | 直接提问「分析这张截图」,模型自动调 `read_clipboard_image` 读系统剪贴板 |
|
||||
| 非多模态 + 想用 Alt+V 粘贴 | 需先在 config.toml 给模型声明 `image_in`(见下),粘贴后直接提问即可——模型看不懂图片内容时会自动调 `read_clipboard_image`(剪贴板仍保留原图) |
|
||||
| 非多模态 + 想用 Alt+V 粘贴 | 纯文本模型需 `/skill kimi-eyes` 配合(见下),单纯粘贴会报错 |
|
||||
| 粘贴被拦截 | 直接告诉模型「我粘贴图片被拦截了」,模型会自动改读剪贴板,无需你保存文件 |
|
||||
|
||||
#### 看图方式 × `image_in` 速查
|
||||
|
||||
| 看图方式 | 需要 `image_in`? | 纯文本可用? | 顺滑度 |
|
||||
| --- | --- | --- | --- |
|
||||
| 截图后随口一问(自动读剪贴板) | 否 | ✅ | 最省事 |
|
||||
| `@路径` / 给路径 | 否 | ✅ | 省事 |
|
||||
| `/skill kimi-eyes` + Alt-V | 是 | ✅ | 多步 |
|
||||
| 单纯 Alt-V | 声明与否 | ❌ | 走不通 |
|
||||
|
||||
> 纯文本模型日常看图优先用前两种;只有「就想用粘贴这个动作」时才走 `/skill kimi-eyes`(见下)。
|
||||
|
||||
#### 想用 Alt+V 粘贴?(给模型声明 `image_in`)
|
||||
|
||||
KimiCode 前端默认会拦截「不支持图片输入」模型的粘贴。在 `~/.kimi-code/config.toml` 给对应模型加上 `image_in` 即可放行:
|
||||
@@ -102,6 +113,37 @@ capabilities = [ "thinking", "tool_use", "image_in" ] # 追加 image_in
|
||||
>
|
||||
> 纯文本模型的正确用法:`@图片路径`(read_image)或截图后直接提问(read_clipboard_image 读剪贴板);粘贴被拦截就直接告诉模型「粘贴被拦截了」,它会自动改读剪贴板。**全程零命令**。
|
||||
|
||||
#### 纯文本模型也想「粘贴看图」?用 `/skill`(推荐)
|
||||
|
||||
上面警告过:纯文本 provider 声明 `image_in` 后走 **Alt-V 粘贴**,图片会作为 `image_url` part 发给 provider 触发 400。但同样声明 `image_in`,改走 **`/skill` 命令**就不会 400——`/skill` 通道把图片渲染成 `Attached image file: <路径>` 的**纯文本路径**,不产生 image part。本插件自带一个 skill 利用这条通道。
|
||||
|
||||
**一次性配置(两步)**:
|
||||
|
||||
1. 给纯文本模型声明 `image_in`(仅为通过 `/skill` 的前端校验;`/skill` 下图片不走 part,**不会**触发 400):
|
||||
|
||||
```toml
|
||||
[models."opencode-go/deepseek-v4-flash"]
|
||||
capabilities = [ "thinking", "tool_use", "image_in" ] # 追加 image_in
|
||||
```
|
||||
|
||||
2. 注册 skill——在 `~/.kimi-code/config.toml` 把插件目录的 `skills/` 加入扫描:
|
||||
|
||||
```toml
|
||||
extra_skill_dirs = [ "D:/AIGC/Plugin/kimi-eyes/skills" ]
|
||||
```
|
||||
|
||||
> 插件装在别处就换成实际的 `kimi-eyes/skills` 目录。重启会话后 `/skill kimi-eyes` 会出现在 `/` 补全菜单。
|
||||
|
||||
**用法**:
|
||||
|
||||
```
|
||||
/skill kimi-eyes 这张图表说明了什么? ← 然后 Alt-V 粘贴图片,回车
|
||||
```
|
||||
|
||||
模型会从消息里的 `Attached image file: <路径>` 提取路径,调用 kimi-eyes 的 `read_image`(外部 VLM 返回文字描述)再作答。全程不产生 image part,纯文本模型不会 400。
|
||||
|
||||
> ⚠️ **声明 `image_in` 后的使用约束**:看图一律走 `/skill kimi-eyes`,**不要**再直接 Alt-V 粘贴(主 prompt 通道仍会把图片作 image part 发给纯文本 provider → 400)。`@图片路径` 和「截图后提问」两种老用法不受影响。
|
||||
|
||||
## 激活与触发
|
||||
|
||||
### 激活(一次性)
|
||||
|
||||
+31
-1
@@ -429,7 +429,10 @@ npx kimi-eyes setup
|
||||
|
||||
<span class="cmt"># 使用</span>
|
||||
@截图.png 这个图里有什么? <span class="cmt"># → read_image</span>
|
||||
分析这张截图 <span class="cmt"># → read_clipboard_image</span></pre>
|
||||
分析这张截图 <span class="cmt"># → read_clipboard_image</span>
|
||||
|
||||
<span class="cmt"># 纯文本模型想「粘贴」看图(需先声明 image_in + 注册 skill,详见 README):</span>
|
||||
/skill kimi-eyes 这张图说明什么? <span class="cmt"># 然后 Alt-V 粘贴图片,回车</span></pre>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
@@ -456,6 +459,33 @@ claude mcp remove kimi-eyes</pre>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<!-- WAYS TO VIEW -->
|
||||
<section class="block" id="ways">
|
||||
<div class="wrap">
|
||||
<h2 class="s-title reveal">看图方式 × <code style="font-family:var(--mono);color:var(--accent)">image_in</code></h2>
|
||||
<p class="s-lead reveal">纯文本主模型下,不同看图姿势对 <code style="font-family:var(--mono)">image_in</code> 的要求和可用性不同。</p>
|
||||
<div class="reveal" style="overflow-x:auto;margin-top:32px">
|
||||
<table style="width:100%;border-collapse:collapse;font-size:14px;min-width:560px">
|
||||
<thead>
|
||||
<tr style="border-bottom:1px solid var(--border)">
|
||||
<th style="text-align:left;padding:12px 10px;font-weight:600;color:var(--text-muted)">看图方式</th>
|
||||
<th style="text-align:left;padding:12px 10px;font-weight:600;color:var(--text-muted)">需要 image_in?</th>
|
||||
<th style="text-align:left;padding:12px 10px;font-weight:600;color:var(--text-muted)">纯文本可用?</th>
|
||||
<th style="text-align:left;padding:12px 10px;font-weight:600;color:var(--text-muted)">顺滑度</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr style="border-bottom:1px solid var(--border-soft)"><td style="padding:12px 10px">截图后随口一问(自动读剪贴板)</td><td style="padding:12px 10px;color:var(--text-dim)">否</td><td style="padding:12px 10px;color:var(--accent)">✅</td><td style="padding:12px 10px;color:var(--text-muted)">最省事</td></tr>
|
||||
<tr style="border-bottom:1px solid var(--border-soft)"><td style="padding:12px 10px"><code style="font-family:var(--mono)">@路径</code> / 给路径</td><td style="padding:12px 10px;color:var(--text-dim)">否</td><td style="padding:12px 10px;color:var(--accent)">✅</td><td style="padding:12px 10px;color:var(--text-muted)">省事</td></tr>
|
||||
<tr style="border-bottom:1px solid var(--border-soft)"><td style="padding:12px 10px"><code style="font-family:var(--mono)">/skill kimi-eyes</code> + Alt-V</td><td style="padding:12px 10px;color:var(--text)">是</td><td style="padding:12px 10px;color:var(--accent)">✅</td><td style="padding:12px 10px;color:var(--text-muted)">多步</td></tr>
|
||||
<tr><td style="padding:12px 10px">单纯 Alt-V</td><td style="padding:12px 10px;color:var(--text-dim)">声明与否</td><td style="padding:12px 10px;color:var(--text-dim)">❌</td><td style="padding:12px 10px;color:var(--text-dim)">走不通</td></tr>
|
||||
</tbody>
|
||||
</table>
|
||||
</div>
|
||||
<p class="s-lead reveal" style="margin-top:20px">日常优先用前两种;只有「就想粘贴」时才走 <code style="font-family:var(--mono)">/skill kimi-eyes</code>。</p>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<!-- TOOLS -->
|
||||
<section class="block" id="tools">
|
||||
<div class="wrap">
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "kimi-eyes",
|
||||
"version": "1.0.6",
|
||||
"version": "1.0.7",
|
||||
"description": "Give non-multimodal models vision: analyze local images and clipboard screenshots through a user-configured VLM API (OpenAI-compatible or Anthropic protocol).",
|
||||
"keywords": [
|
||||
"vision",
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "kimi-eyes",
|
||||
"version": "1.0.6",
|
||||
"version": "1.0.7",
|
||||
"description": "Give non-multimodal models in Kimi Code the ability to analyze images and screenshots via a user-configured VLM (OpenAI-compatible or Anthropic protocol).",
|
||||
"keywords": [
|
||||
"kimi",
|
||||
|
||||
@@ -0,0 +1,36 @@
|
||||
---
|
||||
name: kimi-eyes
|
||||
description: 纯文本主模型的看图入口。通过 /skill kimi-eyes 触发,把粘贴的图片交给 kimi-eyes 配置的外部 VLM 分析,返回文字描述。
|
||||
type: prompt
|
||||
whenToUse: 当用户用 /skill kimi-eyes 发送图片、且当前主模型不具备原生看图能力时
|
||||
---
|
||||
|
||||
# Kimi Eyes — 纯文本主模型看图
|
||||
|
||||
## 为什么需要这个 skill
|
||||
|
||||
`/skill` 通道下,用户粘贴的图片以 `Attached image file: <缓存路径>` 的**纯文本**形式
|
||||
进入对话,不会作为 `image_url` content part 发给主模型。因此**纯文本主模型也不会触发
|
||||
`400 unknown variant image_url` 错误**。本 skill 利用这条文本通道,把图片交给 kimi-eyes
|
||||
已配置的外部 VLM,拿回文字描述后作答。
|
||||
|
||||
## 触发条件
|
||||
|
||||
- **仅当**用户当前消息以 `/skill kimi-eyes` 开头,并附带图片
|
||||
- 用户直接 Alt-V 粘贴图片(没有 `/skill` 前缀)时**不要**走本 skill,按 kimi-eyes 常规机制处理(@图片路径 或 截图后直接提问)
|
||||
- 多模态主模型(已声明 image_in 且 provider 真支持图片)原生看图即可,不需要本 skill
|
||||
|
||||
## 执行步骤
|
||||
|
||||
1. 在用户消息里定位 `Attached image file: <路径>` 文本,提取其中的图片**绝对路径**(缓存目录下的文件)
|
||||
2. 调用 kimi-eyes 的 `read_image` 工具:
|
||||
- `path` = 第 1 步提取到的图片路径
|
||||
- `prompt` = 用户的问题;若用户没给具体问题,使用「详细描述这张图片的内容」
|
||||
3. **不要**使用 `ReadMediaFile`(它对纯文本模型无意义,且会把图片作为 part 回灌,可能再次触发 400)
|
||||
4. `read_image` 返回的是外部 VLM 的**文字描述**,基于这段描述回答用户的问题
|
||||
|
||||
## 失败兜底
|
||||
|
||||
- 消息里没有 `Attached image file:` 路径 → 提示用户:「请输入 `/skill kimi-eyes` 后再粘贴图片」
|
||||
- `read_image` 报「Vision API is not configured」→ 提示用户运行 `npx kimi-eyes setup` 配置视觉 API
|
||||
- `read_image` 报文件不存在 → 提示用户重新粘贴图片(缓存文件可能已失效)
|
||||
Reference in new issue
Block a user