Algorithms on strings, trees, and sequences computer science and computational biology
Dan Gusfield
A comprehensive textbook on string algorithms, suffix trees, and sequence alignment with applications in computational biology.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Dan Gusfield
A comprehensive textbook on string algorithms, suffix trees, and sequence alignment with applications in computational biology.
Franz Josef Och, Hermann Ney
This paper systematically compares statistical and heuristic word alignment models, showing refined models with first-order dependence and fertility significantly outperform simple heuristics.
Omar Shaikh, Michelle Lam, Joey Hejna, et al.
DITTO aligns LLMs to user-specific behaviors using fewer than 10 demonstrations via online imitation learning and preference optimization.
J. Kirchner, Yining Chen, Harri Edwards, et al.
Proposes legibility training via a Prover-Verifier Game to make LLM chain-of-thought reasoning easier for humans to verify, improving trust in model outputs.
Ke Wang, Jiahui Zhu, Minjie Ren, et al.
This survey comprehensively reviews data generation techniques for LLMs, covering data augmentation and synthesis across the entire LLM lifecycle.
A. Sheshadri, John Hughes, Julian Michael, et al.
This paper analyzes 25 language models to understand why only 5 exhibit alignment faking, finding that post-training variations in refusal behavior largely explain differences.
Unknown
QALIGN uses MCMC sampling and Minimum Bayes Risk to align language model outputs at test time without retraining.
Unknown
RAFT iteratively fine-tunes generative models on top-ranked samples to align them with a reward function, improving stability and efficiency over RLHF.
Unknown
SweEval is a cross-lingual safety benchmark that evaluates LLMs on handling swear words across 8 languages, revealing higher vulnerability in Indic languages and transliterated contexts.
Daniel Bolya, Po-Yao Huang, Peize Sun, et al.
Perception Encoder is a vision encoder trained via contrastive vision-language learning that achieves state-of-the-art results across diverse tasks by extracting strong general embeddings from intermediate layers.
Unknown
LLMLingua introduces a coarse-to-fine prompt compression method achieving up to 20x compression with minimal performance loss.
Unknown
mmE5 is a multimodal multilingual embedding model trained on synthetic data via a framework that ensures broad task/language coverage, robust cross-modal alignment, and high fidelity through self-evaluation and refinement.