Tool: computer — Direct Vision ควบคุมจอจริงแบบ screenshot → action → re-observe

ภาพประกอบ: Tool: computer — Direct Vision ควบคุมจอจริงแบบ screenshot → action → re-observe

computer ควบคุมเมาส์/คีย์บอร์ดจริงของเครื่องแทนคน — click, type, key, scroll, drag, เปิดแอปและเปิด URL ทีละ action แล้วดูผลลัพธ์ใหม่ก่อนตัดสินใจทำขั้นถัดไป

รุ่นปัจจุบันเป็น capability-aware: ถ้า main model รองรับ vision computer จะใช้ direct screenshot vision โดยไม่มี inner vision model แยกซ่อนอยู่ใน tool. ค่า default ของ Agent TH ยังเป็น Qwen3-14B แบบ text-only สำหรับเครื่อง 24GB+ แต่ Mac RAM 16GB สามารถเลือก Qwen3.5-9B-4bit VLM เพื่อใช้ computer แบบ direct vision ได้

Flow หลัก: screenshot → aim → action → re-observe

capture screenshot
      ↓
ส่งภาพให้ main VLM + [OBS] จาก AX/OCR
      ↓
Agent เลือก target/action
      ↓
computer ทำ action จริง
      ↓
capture/observe ใหม่
      ↓
ตรวจ visible effect
      ↓
สำเร็จค่อยทำต่อ / ไม่เปลี่ยนให้ recover

จุดสำคัญคือหลัง action ที่เปลี่ยน UI ระบบไม่ควรใช้ภาพ/observation เก่าต่อแล้ว “คิดเอาเอง” ว่าคลิกสำเร็จ แต่ต้องดู state ใหม่

เมื่อเลือก vision-capable model: Main VLM เป็นคนมองภาพ

เมื่อใช้ Qwen3.5-9B VLM หรือ Qwen3.6-35B ที่รองรับ image input screenshot ล่าสุดจะถูก publish ให้ main agent เห็น pixels โดยตรง

ทำให้ Agent ใช้ข้อมูลที่ OCR อย่างเดียวมองไม่เห็นได้ เช่น:

  • ตำแหน่งและ layout
  • ไอคอน
  • สี/สถานะ visual
  • dialog ที่องค์ประกอบไม่ได้มี text ครบ
  • ความสัมพันธ์เชิงพื้นที่ระหว่าง element

ไม่มีการเรียกโมเดล vision ตัวเล็กอีกตัวเพื่อ “บรรยายภาพให้ 35B ฟัง”

Accessibility + OCR ยังมีไว้ทำไม

direct vision ไม่ได้แปลว่า structured sensors ไม่มีประโยชน์

computer รวม observation จาก:

  • macOS Accessibility (AX) — role, label, focus, element ที่ interact ได้
  • Apple Vision OCR — text บนหน้าจอ
  • visual screenshot — pixels จริงที่ main VLM เห็น

ข้อมูล AX/OCR ถูกแนบเป็น [OBS] เพื่อให้ Agent อ้าง element/text ได้แม่นขึ้น และช่วยตรวจ state หลัง action

ถ้า Accessibility permission ยังไม่เปิด tool สามารถรายงาน ax=permission_required โดยไม่ crash; visual path ยังทำงานตาม capability ของ model

Qwen3-14B default: fail closed ก่อนแตะ desktop

computer ต้องใช้ vision-capable model ดังนั้น Qwen3-14B default ยังใช้ Agent TH ด้าน file/code/web/RAG/MCP/tool calling ได้ แต่ computer จะไม่พยายามขับจอจากข้อมูลที่มองไม่เห็น. หากเป็น Mac RAM 16GB ให้เลือก Qwen3.5-9B VLM เพื่อเปิด direct screenshot vision

ก่อน screenshot/action path runtime จะทำ capability probe แบบไม่ mutate desktop ถ้าพบว่า model เป็น text-only — เช่น Qwen3-14B ใน configuration ที่ทดสอบกับโปรเจกต์นี้ — tool จะคืน:

[unsupported] computer requires a vision-capable model

และหยุดก่อน:

  • ถ่าย screenshot
  • click
  • type
  • key
  • scroll
  • เปิด app/URL

ไม่มี OCR fallback เพื่อให้ text-only model “ขับจอแบบเดา ๆ” เพราะความเสี่ยงสูงเกินไป

Observation มี owner และ lifecycle ของตัวเอง

state ของ computer แยกจาก read_image

  • screenshot ของ desktop เป็น computer-owned state
  • observation มี id/version
  • inspect สามารถ publish crop/detail ใหม่ของ desktop state
  • หลัง mutation ต้อง refresh observation
  • state ถูก reset ตาม turn/lifecycle

จึงไม่เอาภาพจากไฟล์ทั่วไปหรือภาพเก่ามาปะปนกับสถานะ desktop ล่าสุด

Silent failure recovery

action บน UI อาจไม่เกิดผลแม้ mouse click จะถูกส่งออกไปจริง เช่น target ขยับ, focus ไม่ถูก, overlay บัง หรือ app ยังโหลดไม่เสร็จ

ดังนั้น computer path ตรวจทั้ง visual/AX/OCR effect หลัง action ถ้าไม่เห็นการเปลี่ยนแปลง:

  1. ไม่ claim success
  2. observe ใหม่
  3. หา target ปัจจุบันอีกครั้ง
  4. retry ด้วย element/ตำแหน่งใหม่เมื่อเหมาะสม

นี่สำคัญกว่าการพยายามกดซ้ำ coordinate เดิมโดยไม่ดูหน้าจอ

Action ที่รองรับ

กลุ่มหลัก ได้แก่:

  • see, inspect
  • click, double_click, triple_click, right_click
  • hover, drag
  • type
  • key
  • scroll
  • open_app
  • open_url

บาง action รองรับ expect เพื่อให้ tool รอดูเงื่อนไขหลัง action เช่น app/window/text/focus แล้วรายงานว่า expectation สำเร็จหรือไม่

Guard ที่ทำงานด้วย code

เพราะ computer ทำ action จริงบนเครื่อง safety ไม่ควรฝากไว้กับ prompt อย่างเดียว

runtime มี guard เช่น:

  • destructive-looking click/drag/hotkey ต้องสอดคล้องกับ user intent ปัจจุบัน
  • action limit ต่อ turn
  • turn ที่ยิงจาก awake เข้มกว่า interactive turn และบล็อก destructive action
  • state reset ทุก turn
  • capability gate มาก่อน desktop side effect
  • secure-field verification ไม่เก็บเนื้อหาลับกลับมาเป็น state

โมเดลเป็นคนเสนอ action แต่ runtime เป็นคนตัดสินว่าสามารถ execute ได้หรือไม่

ต่างจาก read_image

read_image มีหน้าที่ “ดูภาพ” ส่วน computer มีหน้าที่ “ดู desktop แล้วลงมือ”

ทั้งคู่ใช้ main-model direct vision แต่มี state/lifecycle แยกกัน และ computer มี safety/effect verification เพิ่มเพราะ action เปลี่ยน state จริงของเครื่อง

ทำไมเรื่องนี้สำคัญ

การเปลี่ยนจาก OCR-only → direct screenshot vision ทำให้ Agent เข้าใจ desktop ได้มากขึ้นโดยไม่เพิ่ม inner agent หรือ inner VLM อีกชั้น ขณะเดียวกัน capability gate และ deterministic guard ทำให้ text-only model ไม่ถูกปล่อยให้ควบคุม desktop อย่างไม่ปลอดภัย

แนวคิดจึงเป็น:

ให้ model เห็นมากขึ้น แต่ให้ code คุมสิทธิ์มากขึ้นด้วย

อ่านเพิ่มเติม

ความคิดเห็น

กำลังโหลดความคิดเห็น...