All Documents
3,528 documents available
Smart Hybrid Retrieval - Implementation Summary
Describes a 4-phase retrieval algorithm combining semantic search, graph expansion, completeness checks, and multi-factor ranking for knowledge graphs.
Property chunking
Defines a YAML schema and configuration pattern for splitting large property lists into smaller API requests and merging results back into complete records.
Long Context Chunking
Splits long documents into overlapping chunks and averages their embeddings to avoid context limit errors.
MetaAST-Enhanced Retrieval
Describes how MetaAST metadata boosts code search ranking, enables cross-language construct lookup, and expands queries with synonyms.
chunking
Explains the chunking architecture in Next.js Turbopack, covering graph walk, module batching, merging, chunk splitting, and output generation.
ChonkyDB One-Shot Retrieval Target
Proposes a new ChonkyDB endpoint that returns final structured records in one request, eliminating separate hydration roundtrips for Twinr's low-latency retrieval.
Search and Retrieval
Documents a 6-stage RAG search pipeline with hybrid retrieval, reranking, Chain of Thought reasoning, and source attribution.
Marker Page Chunking Benchmark
Benchmarks Marker chunk sizes (10, 20, 30) for PDF parsing speed and GPU memory, recommending chunk size 10 for production.
Better File Chunking
Opens a discussion on chunking strategies for IPFS, aiming to reduce algorithm proliferation and improve default chunking.
prompt-context-retrieval
title: Just-in-Time Context Retrieval
curl-content-retrieval
Fetches a specific URL with curl to verify web resource availability and content for subdomain takeover validation.
Error Context Retrieval
Retrieves stored error patterns before build/test commands execute, letting Claude anticipate and avoid known failures.
Bitswap Retrieval
Explains how to configure and run booster-bitswap alongside boostd to serve Filecoin retrievals over the Bitswap protocol in two modes.
Retrieval
Explains synchronous and asynchronous retrieval from a SQLite database using DBFlow's query language and transaction system.
Text Retrieval Benchmark
Runs a text retrieval benchmark using BEIR datasets, measuring recall, precision, NDCG, and MRR without LLM generation.
Lab 3: Information Retrieval (Week 4)
Walks student groups through computing tf-idf cosine similarity, comparing it to dense retrieval, and calculating precision and recall.
Data Retrieval
Defines Python functions to retrieve and clean UK electricity generation, price, temperature and emissions data from the Electric Insights API.
Retrieving data from VLM's pretraining dataset
Walks through a four-step pipeline to retrieve and sample 500 images per class from LAION-400M using synonym matching and T2T ranking.
Audio Retrieval with Natural Language Queries
Implements audio retrieval with natural language queries using collaborative experts, providing training and evaluation scripts for multiple datasets.
PRD: Historical Retrieval System
Defines a system that records past agent actions and injects relevant successes and failures into new sessions to avoid repeated mistakes.
Retrieval Evaluation for fasten-answers-ai
Walks through evaluating retrieval accuracy, MRR, precision, and recall using Elasticsearch and a Q&A dataset.
Information Retrieval - Topic Modelling Relevent Part
Serves as a structured reference for information retrieval concepts, covering models, term weighting, evaluation measures, and link analysis.
Data Retrieval
Downloads UK electricity generation and pricing data from Electric Insights and Energy Charts APIs using the moepy library, then visualises fuel-mix time-series.
RETRIEVAL.md — Memory Scan Protocol
Defines a five-step internal memory retrieval protocol that surfaces results only in LiveHud gauges, never as visible logs.