Tool: computer — Direct Vision ควบคุมจอจริงแบบ screenshot → action → re-observe

computer ควบคุมเมาส์/คีย์บอร์ดจริงของเครื่องแทนคน — click, type, key, scroll, drag, เปิดแอปและเปิด URL ทีละ action แล้วดูผลลัพธ์ใหม่ก่อนตัดสินใจทำขั้นถัดไป
รุ่นปัจจุบันเป็น capability-aware: ถ้า main model รองรับ vision computer จะใช้ direct screenshot vision โดยไม่มี inner vision model แยกซ่อนอยู่ใน tool. ค่า default ของ Agent TH ยังเป็น Qwen3-14B แบบ text-only สำหรับเครื่อง 24GB+ แต่ Mac RAM 16GB สามารถเลือก Qwen3.5-9B-4bit VLM เพื่อใช้ computer แบบ direct vision ได้
Flow หลัก: screenshot → aim → action → re-observe
capture screenshot
↓
ส่งภาพให้ main VLM + [OBS] จาก AX/OCR
↓
Agent เลือก target/action
↓
computer ทำ action จริง
↓
capture/observe ใหม่
↓
ตรวจ visible effect
↓
สำเร็จค่อยทำต่อ / ไม่เปลี่ยนให้ recover
จุดสำคัญคือหลัง action ที่เปลี่ยน UI ระบบไม่ควรใช้ภาพ/observation เก่าต่อแล้ว “คิดเอาเอง” ว่าคลิกสำเร็จ แต่ต้องดู state ใหม่
เมื่อเลือก vision-capable model: Main VLM เป็นคนมองภาพ
เมื่อใช้ Qwen3.5-9B VLM หรือ Qwen3.6-35B ที่รองรับ image input screenshot ล่าสุดจะถูก publish ให้ main agent เห็น pixels โดยตรง
ทำให้ Agent ใช้ข้อมูลที่ OCR อย่างเดียวมองไม่เห็นได้ เช่น:
- ตำแหน่งและ layout
- ไอคอน
- สี/สถานะ visual
- dialog ที่องค์ประกอบไม่ได้มี text ครบ
- ความสัมพันธ์เชิงพื้นที่ระหว่าง element
ไม่มีการเรียกโมเดล vision ตัวเล็กอีกตัวเพื่อ “บรรยายภาพให้ 35B ฟัง”
Accessibility + OCR ยังมีไว้ทำไม
direct vision ไม่ได้แปลว่า structured sensors ไม่มีประโยชน์
computer รวม observation จาก:
- macOS Accessibility (AX) — role, label, focus, element ที่ interact ได้
- Apple Vision OCR — text บนหน้าจอ
- visual screenshot — pixels จริงที่ main VLM เห็น
ข้อมูล AX/OCR ถูกแนบเป็น [OBS] เพื่อให้ Agent อ้าง element/text ได้แม่นขึ้น และช่วยตรวจ state หลัง action
ถ้า Accessibility permission ยังไม่เปิด tool สามารถรายงาน ax=permission_required โดยไม่ crash; visual path ยังทำงานตาม capability ของ model
Qwen3-14B default: fail closed ก่อนแตะ desktop
computer ต้องใช้ vision-capable model ดังนั้น Qwen3-14B default ยังใช้ Agent TH ด้าน file/code/web/RAG/MCP/tool calling ได้ แต่ computer จะไม่พยายามขับจอจากข้อมูลที่มองไม่เห็น. หากเป็น Mac RAM 16GB ให้เลือก Qwen3.5-9B VLM เพื่อเปิด direct screenshot vision
ก่อน screenshot/action path runtime จะทำ capability probe แบบไม่ mutate desktop ถ้าพบว่า model เป็น text-only — เช่น Qwen3-14B ใน configuration ที่ทดสอบกับโปรเจกต์นี้ — tool จะคืน:
[unsupported] computer requires a vision-capable model
และหยุดก่อน:
- ถ่าย screenshot
- click
- type
- key
- scroll
- เปิด app/URL
ไม่มี OCR fallback เพื่อให้ text-only model “ขับจอแบบเดา ๆ” เพราะความเสี่ยงสูงเกินไป
Observation มี owner และ lifecycle ของตัวเอง
state ของ computer แยกจาก read_image
- screenshot ของ desktop เป็น computer-owned state
- observation มี id/version
- inspect สามารถ publish crop/detail ใหม่ของ desktop state
- หลัง mutation ต้อง refresh observation
- state ถูก reset ตาม turn/lifecycle
จึงไม่เอาภาพจากไฟล์ทั่วไปหรือภาพเก่ามาปะปนกับสถานะ desktop ล่าสุด
Silent failure recovery
action บน UI อาจไม่เกิดผลแม้ mouse click จะถูกส่งออกไปจริง เช่น target ขยับ, focus ไม่ถูก, overlay บัง หรือ app ยังโหลดไม่เสร็จ
ดังนั้น computer path ตรวจทั้ง visual/AX/OCR effect หลัง action ถ้าไม่เห็นการเปลี่ยนแปลง:
- ไม่ claim success
- observe ใหม่
- หา target ปัจจุบันอีกครั้ง
- retry ด้วย element/ตำแหน่งใหม่เมื่อเหมาะสม
นี่สำคัญกว่าการพยายามกดซ้ำ coordinate เดิมโดยไม่ดูหน้าจอ
Action ที่รองรับ
กลุ่มหลัก ได้แก่:
see,inspectclick,double_click,triple_click,right_clickhover,dragtypekeyscrollopen_appopen_url
บาง action รองรับ expect เพื่อให้ tool รอดูเงื่อนไขหลัง action เช่น app/window/text/focus แล้วรายงานว่า expectation สำเร็จหรือไม่
Guard ที่ทำงานด้วย code
เพราะ computer ทำ action จริงบนเครื่อง safety ไม่ควรฝากไว้กับ prompt อย่างเดียว
runtime มี guard เช่น:
- destructive-looking click/drag/hotkey ต้องสอดคล้องกับ user intent ปัจจุบัน
- action limit ต่อ turn
- turn ที่ยิงจาก
awakeเข้มกว่า interactive turn และบล็อก destructive action - state reset ทุก turn
- capability gate มาก่อน desktop side effect
- secure-field verification ไม่เก็บเนื้อหาลับกลับมาเป็น state
โมเดลเป็นคนเสนอ action แต่ runtime เป็นคนตัดสินว่าสามารถ execute ได้หรือไม่
ต่างจาก read_image
read_image มีหน้าที่ “ดูภาพ” ส่วน computer มีหน้าที่ “ดู desktop แล้วลงมือ”
ทั้งคู่ใช้ main-model direct vision แต่มี state/lifecycle แยกกัน และ computer มี safety/effect verification เพิ่มเพราะ action เปลี่ยน state จริงของเครื่อง
ทำไมเรื่องนี้สำคัญ
การเปลี่ยนจาก OCR-only → direct screenshot vision ทำให้ Agent เข้าใจ desktop ได้มากขึ้นโดยไม่เพิ่ม inner agent หรือ inner VLM อีกชั้น ขณะเดียวกัน capability gate และ deterministic guard ทำให้ text-only model ไม่ถูกปล่อยให้ควบคุม desktop อย่างไม่ปลอดภัย
แนวคิดจึงเป็น:
ให้ model เห็นมากขึ้น แต่ให้ code คุมสิทธิ์มากขึ้นด้วย
ความคิดเห็น
กำลังโหลดความคิดเห็น...