Preprint2025
ThinkPRM
Unknown
A generative PRM trained on small synthetic data outperforms discriminative PRMs and LLM-as-a-Judge baselines in step-by-step reasoning verification.
0Apr 1, 2025Reasoning
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
A generative PRM trained on small synthetic data outperforms discriminative PRMs and LLM-as-a-Judge baselines in step-by-step reasoning verification.
Unknown
A pipeline that generates system messages and aligned assistant responses for SFT datasets lacking them, using LLM-as-a-judge verification.
Unknown
Proposes an iterative self-training method using only synthetic data to improve LLM-as-a-Judge without human annotations.