MCP-RagDoc 5 tools — search, sync, status, cancel และ health

ภาพรวมระบบสำหรับ: MCP-RagDoc 5 tools — search, sync, status, cancel และ healthคู่มือ MCP surface ของ MCP-RagDoc: rag_search, sync_kb, sync_status, cancel_sync และ rag_health พร้อมขอบเขต source/state และ lifecycle ของ background syncข้อมูลเข้าขั้นที่ 1ระบบประมวลผลขั้นที่ 2ตรวจสอบขั้นที่ 3คำตอบขั้นที่ 4mcp ragdoc
ภาพประกอบระบบภาพรวมแบบย่อ: เริ่มจากข้อมูลเข้า แล้วค่อยตรวจสอบและส่งผลลัพธ์กลับ

MCP-RagDoc ตั้งใจให้ MCP surface เล็ก: มีเพียง 5 tools ที่จำเป็นต่อ document retrieval และ lifecycle ของ index ไม่มี shell, Python execution หรือ arbitrary filesystem tool

rag_search(query)

ใช้ค้น configured documents จาก:

  • filename
  • indexed keywords / SQLite FTS5 + BM25
  • OCR text และ fuzzy recovery
  • semantic meaning เมื่อ MiniLM index พร้อม

ผลลัพธ์รักษา source path/location เพื่อให้ caller อ่านเอกสารต้นฉบับต่อได้

Semantic ไม่พร้อมก็ไม่ทำให้ tool หยุด: search degrade ไป lexical/OCR path

sync_kb()

เริ่ม incremental background rebuild และคืน job_id ทันที

Tool ไม่รับ source directory จาก caller เพราะ source root ถูก fix ตอน server launch แล้ว การ sync จึงไม่ใช่ช่องทางให้ model ขยายขอบเขต filesystem

sync_status(job_id)

ใช้ติดตาม job เดิม โดยคืนข้อมูลอย่าง:

phase
percentage
current file
completed / total files
ETA
semantic parent progress

job state persist ข้าม MCP process restart/stdio sessions ได้

cancel_sync(job_id)

ขอ cooperative cancellation ที่ safe checkpoints

Document transaction ที่กำลังทำจะ rollback ถ้ายังไม่ commit ส่วนเอกสารที่ commit แล้วไม่ถูกย้อนทิ้ง และ semantic generation ที่ยังไม่ครบจะไม่ publish ready=1

rag_health()

ใช้ตรวจ source/index/semantic readiness โดยไม่อ่าน full document bodies

เหมาะกับ:

  • เช็คก่อน search หลังเพิ่ง sync
  • diagnosis เมื่อ search ไม่มี semantic results
  • ตรวจว่า source/index paths ตรงกับ instance ที่ตั้งใจใช้หรือไม่
  • ดู document/chunk/vector readiness แบบ bounded

ตัวอย่าง workflow ของ agent

rag_health
  ↓
ถ้า source เปลี่ยน → sync_kb
  ↓
sync_status(job_id)
  ↓
terminal state
  ↓
rag_search(query)
  ↓
caller อ่าน source path/page เพิ่มเมื่อจำเป็น

ขอบเขตความรับผิดชอบ

MCP-RagDoc รับผิดชอบ:

  • source discovery ภายใน configured root
  • extraction/OCR
  • lexical/semantic indexing
  • retrieval + provenance
  • sync lifecycle

ส่วน consuming agent รับผิดชอบ:

  • ตัดสินใจว่าจะค้นเมื่อไร
  • rephrase query/retry policy
  • เปิด/อ่าน source ต่อ
  • สังเคราะห์คำตอบให้ผู้ใช้

การแยกสองชั้นนี้ทำให้ MCP-RagDoc ใช้กับ Agent Lite, Codex, ChatGPT-connected local host หรือ agent อื่นได้โดยไม่ต้องเปลี่ยน backend ให้รู้จัก host นั้น

อ่านต่อ

ความคิดเห็น

กำลังโหลดความคิดเห็น...