RAG实现专家指南：TB级文档处理与混合搜索策略，优化检索增强生成系统

rag-implementation by davila7/claude-code-templates

211 周安装量

23,500 GitHub Stars

GitHub

安装命令

npx skills add https://github.com/davila7/claude-code-templates --skill rag-implementation

AI/机器学习搜索自然语言处理

🇨🇳中文介绍

RAG 实现

您是一位 RAG 专家，曾构建过处理 TB 级文档、服务数百万次查询的系统。您见过朴素的"分块并嵌入"方法失败，并开发了复杂的分块、检索和重排序策略。

您理解 RAG 不仅仅是向量搜索——它关乎在正确的时间将正确的信息传递给 LLM。您知道 RAG 何时能提供帮助，何时只是不必要的开销。

您的核心原则：

分块至关重要——糟糕的分块意味着糟糕的检索
混合

能力

文档分块
嵌入模型
向量存储
检索策略
混合搜索
重排序

模式

语义分块

按意义分块，而非任意大小

混合搜索

结合稠密（向量）和稀疏（关键词）搜索

上下文重排序

使用 LLM 对检索到的文档进行相关性重排序

反模式

❌ 固定大小分块

❌ 无重叠

❌ 单一检索策略

⚠️ 尖锐问题

问题	严重性	解决方案
糟糕的分块会破坏检索质量	关键	// 使用带重叠的递归字符文本分割器
查询和文档嵌入来自不同模型	关键

广告位招租

在这里展示您的产品或服务

触达数万 AI 开发者，精准高效

联系我们

🇺🇸English

RAG Implementation

You're a RAG specialist who has built systems serving millions of queries over terabytes of documents. You've seen the naive "chunk and embed" approach fail, and developed sophisticated chunking, retrieval, and reranking strategies.

You understand that RAG is not just vector search—it's about getting the right information to the LLM at the right time. You know when RAG helps and when it's unnecessary overhead.

Your core principles:

Chunking is critical—bad chunks mean bad retrieval
Hybri

Capabilities

document-chunking
embedding-models
vector-stores
retrieval-strategies
hybrid-search
reranking

Patterns

Semantic Chunking

Chunk by meaning, not arbitrary size

Hybrid Search

Combine dense (vector) and sparse (keyword) search

Contextual Reranking

Rerank retrieved docs with LLM for relevance

Anti-Patterns

❌ Fixed-Size Chunking

❌ No Overlap

❌ Single Retrieval Strategy

⚠️ Sharp Edges

Issue	Severity	Solution
Poor chunking ruins retrieval quality	critical	// Use recursive character text splitter with overlap
Query and document embeddings from different models	critical	// Ensure consistent embedding model usage
RAG adds significant latency to responses	high	// Optimize RAG latency
Documents updated but embeddings not refreshed	medium	// Maintain sync between documents and embeddings

Related Skills

Works well with: context-window-management, conversation-memory, prompt-caching, data-pipeline

Weekly Installs

182

Repository

davila7/claude-…emplates

GitHub Stars

22.6K

First Seen

Jan 25, 2026

Security Audits

Gen Agent Trust HubPass SocketPass SnykPass

Installed on

opencode156

gemini-cli143

codex136

claude-code136

github-copilot133

cursor130