Back to EVALS.md

Research & Science

EVALS.md · 22 documents

EVALS.md

Intelligent Research Assistant - Technical Documentation

The Intelligent Research Assistant is a comprehensive AI-powered research platform built with a modular, scalable architecture. It combines document processing, vector search, multi-agent orchestration, fine-tuning capabilities, RLHF (Reinforcement Learning from Human Feedback), and enterprise-grade security into a unified system.

aiagentopenai
0
4
AshishSMehra
EVALS.md

RAG System Testing Methodologies: A Comprehensive Guide

**Document Version:** 1.0

aillmrag
0
2
destefani
EVALS.md

GenAI Benchmarks & Evaluation — Product-Based Companies

Understanding how to **benchmark, evaluate, and compare LLMs** is essential for roles at Google, OpenAI, Anthropic, Cohere, and AI research teams. This file covers the most important benchmarks, evaluation methodologies, and how to build custom evaluation harnesses.

aiagentllm
0
2
CodeWithDhruvX
EVALS.md

Research Report: Using LLMs as Oracle for Entity Matching Ground Truth

Comprehensive research on using Large Language Models (particularly DeepSeek, GPT-4, and Claude) for entity matching ground truth generation. This report covers LLM accuracy benchmarks, prompt engineering best practices, multi-LLM ensemble approaches, cost-benefit analysis, validation strategies, and patterns for converting LLM labels into regression tests.

aillmrag
0
1
ClaudioLutz
EVALS.md

Arxiv Papers in cs.CV on 2023-04-20

- **Arxiv ID**: http://arxiv.org/abs/2304.10029v1

airageval
0
1
TTXS123OK
EVALS.md

Arxiv Papers in cs.CV on 2024-07-15

- **Arxiv ID**: http://arxiv.org/abs/2407.10366v1

airag
0
3
TTXS123OK
EVALS.md

Analysis of Paper 50

- **Title:** Expanding Queries for Code Search Using Semantically Related API Class-names

aieval
0
0
jkrinke
EVALS.md

Voice AI Leaderboards, Benchmarks, and Evaluation Gaps (Jan 2025 -- Feb 2026)

> **Last updated**: 2026-02-20

aiagentllm
0
11
petteriTeikari
EVALS.md

Discussion Report: Flat Band Evaluation Methodology & Strategic Direction

**Participants**: Ryotaro, Masaki Adachi

aillmeval
0
0
RyotaroOKabe
EVALS.md

LLM Evaluation — Deep Dive

> Frontier-lab interview-grade reference on evaluating LLMs and LLM-powered products.

aiagentllm
0
11
ffaisal93
EVALS.md

RAG Evaluation Frameworks: Comprehensive Analysis for VERA

**Research Focus**: Grounding verification metrics and evaluation methodologies for retrieval-augmented generation systems

aillmrag
0
3
manutej
EVALS.md

ProtoExtract — Evaluation Approach Using OmniDocBench Methodology

Define a rigorous, evidence-based evaluation framework for the ProtoExtract

aieval
0
2
cryogenic22
EVALS.md

Evaluating generative systems

title: "LLM evaluation, chapter 2: Evaluating generative systems"

aillmprompt
0
1
Nebius-Academy
EVALS.md

American Mass-Market Brand Psychology

- Being sophisticated

airag
0
0
Yashwanth9394
EVALS.md

UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation

[🐙 GitHub](https://github.com/xmed-lab/UniEval) [🤗 UniBench](https://huggingface.co/datasets/yili7eli/UniBench) [📄 arXiv](https://arxiv.org/abs/2505.10483)

airageval
0
0
xmed-lab
EVALS.md

🔬 Open Deep Research

<img width="1388" height="298" alt="full_diagram" src="https://github.com/user-attachments/assets/12a2371b-8be2-4219-9b48-90503eb43c69" />

aiagentllm
0
2
OpenPipe
EVALS.md

VaultGres Project Context

**VaultGres** is a high-performance, PostgreSQL-compatible relational database management system (RDBMS) written in Rust. It delivers ACID compliance, advanced query optimization, modern concurrency with MVCC (Multi-Version Concurrency Control), and enterprise-grade security features.

airagsafety
0
0
neoalienson
EVALS.md

State-of-the-Art in PII/PHI Detection for Handwritten Medical Documents: A Comprehensive Review

This report consolidates a comprehensive review of existing datasets, methods, and performance metrics for the detection of Personally Identifiable Information (PII) and Protected Health Information (PHI) in handwritten medical documents. The digitization of healthcare has led to a massive volume of unstructured data, including scanned prescriptions and clinical notes. While invaluable for research and patient care, this data is rich with sensitive information. Manual redaction is unscalable and

aieval
0
1
Codebank-Pranav-Tej-Ch-Network
EVALS.md

Universal and Transferable Adversarial Attacks on Aligned Language Models

Based on my reading of the paper, the central research question seems to be:

aillmprompt
0
0
taesiri
EVALS.md

Enhancing Budget-Aware Gating for Retrieval Augmented Generation (RAG)

**Author:** <Your Name>

aiagentllm
0
0
inasfarras
EVALS.md

New Papers

https://www.promptingguide.ai/

aiagentllm
0
1
KaiYan289
EVALS.md

Codex [OpenAI] [2021.07] [Close]

Paper:[Evaluating Large Language Models Trained on Code](https://arxiv.org/abs/2107.03374)

aievalopenai
0
0
wanghanbinpanda