Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenizatio
Visual tokenization method using multi-layer representation fusion from frozen vision encoders for better reconstruction.
Visual tokenization method using multi-layer representation fusion from frozen vision encoders for better reconstruction.
Study of involuntary information leakage from prompted secrets in language model outputs using multi-model detection.
Identifies format confound in chain-of-thought corruption studies that detects answer text location rather than actual computation.
Benchmark suite for evaluating physical reasoning and dynamics accuracy in generative video world models.
Empirical evaluation of domain-adapted language models versus general-purpose LLMs for cybersecurity threat modeling tasks.
Policy gradient method analysis for reinforcement learning in non-Markovian environments with internal state management.
Algebraic framework for latent action models in vision-language-action robotics using action-free video data.
Sparse latent steering technique for interpretable control of molecular editing properties in LLMs with explicit property handles.
Multi-phase pretraining framework for learning representations from electronic health records using joint-embedding predictive architecture.
Training-free method for aligning LLM behavior across cultures without fine-tuning or model access, using black-box prompting techniques.
Pi-Serini search agent system pairing lexical retrieval with frontier LLMs for agentic research tasks, evaluating if BM25 suffices for modern reasoning-capable agents.
Multimodal benchmark dataset with 18,000 samples for evaluating AI-assisted CAD program generation from images and 3D observations.
Benchmark for evaluating LLMs and agents on virtual cell modeling tasks, testing in silico phenotypic screening and biological discovery prediction.
Formal framework for shielding techniques in probabilistic Markov decision processes, extending safety guarantees for autonomous agents with acceptable failure probabilities.
Analysis of on-policy distillation for training reasoning models, investigating when teacher-student supervision helps or hurts performance on token-level tasks.
Study of autonomous data engineering for ML systems, automating dataset discovery, adaptation, and validation to reduce manual data engineering workflows.
Research on engineering robustness into AI agents by applying traditional software engineering processes like testing, adversarial evaluation, and staged deployment instead of on-the-fly synthesis.
Continuous diffusion language models using minimal adaptation to match effectiveness of leading discrete-token language model approaches.
PolyMATH benchmark with 5,000 images evaluating multimodal LLM visual comprehension and abstract reasoning across 10 cognitive challenge categories.
DSGBench evaluation platform for LLM-based agents in strategic games assessing long-horizon reasoning, multi-agent interaction, and decision-making.
Framework using LLMs to automate energy-aware refactoring of parallel scientific code focusing on energy efficiency beyond execution time.
LLM-augmented retrosynthesis system for chemical synthesis and drug development combining ML and LLMs to navigate combinatorial pathway space.
Scalable Bayesian planner for multimodal theory-of-mind reasoning inferring beliefs and intentions without task-specific priors.
Planning-based framework for efficient LLM collaboration combining large and small models to reduce inference costs while maintaining performance.
HAMLET framework combines hierarchical multi-agent LLMs for interactive theatrical experiences with embodied interaction and initiative.
Study of algorithmic recourse for ML-driven decisions addressing multi-stakeholder scenarios with shared constraints.
Systematic comparison of reasoning vs non-reasoning LLMs in judge role evaluating accuracy, efficiency, and robustness on small models.
Empirical scaling laws for language model merging showing power law relationship between model size, expert number, and merging performance.
CritPt benchmark evaluates LLM reasoning on complex open-ended frontier physics research challenges beyond high-school math and coding.
Framework using situational judgment tests and multidimensional item response theory to measure stable behavioral tendencies in persona-conditioned LLMs.
Consensus sampling algorithm aggregates multiple probability distributions to improve generative AI safety with architecture-agnostic approach.
Analysis of neural complex query answering over knowledge graphs comparing learned patterns with training-free query relaxation strategies.
Benchmark for evaluating outcome-driven constraint violations in autonomous AI agents, addressing safety and alignment in high-stakes deployment scenarios.
Recursive Language Models enable LLMs to process arbitrarily long prompts through inference-time scaling via recursive self-calling over prompt snippets.
Batch-of-Thought method processes related queries jointly to improve LLM reasoning by identifying high-quality reasoning templates and detecting errors through consistency analysis.
Framework for characterizing and measuring homogenization and mode collapse in LLMs as an AI safety concern affecting diversity.
Study on inter-rater reliability limitations of human feedback for mental health LLM evaluation, questioning assumptions in RLHF approaches.
Method using concise geometric descriptions to improve multimodal LLM performance on plane geometry problem solving tasks.
Proposes Controllable Information Production (CIP) as a principled intrinsic motivation framework for training autonomous agents without external rewards.
Research benchmark (SayNext-Bench) examining why LLMs struggle with next-utterance anticipation in dialogue compared to human multimodal understanding.
M2CL framework improves multi-agent discussion by learning context representations that align individual LLM instances toward coherent solutions.
TodyComm enables dynamic communication topology for multi-agent LLM systems that adapts across conversation rounds based on task progression.
Agent-Omit adaptively omits unnecessary context during multi-turn agent interactions to improve efficiency without sacrificing performance.
Decision-theoretic framework validates whether LLMs hold coherent beliefs by comparing probability judgments to decision behavior.
HyPER framework optimizes test-time compute for LLM reasoning by balancing exploration-exploitation through hypothesis path expansion and reduction.
REVIS training-free framework reduces object hallucination in vision-language models through sparse latent steering.
VeRO evaluation harness systematically measures coding agent performance on agent optimization through iterative edit-execute-evaluate cycles.
Metacognitive behavioral tuning improves LLM multi-hop reasoning by strengthening self-regulation of intermediate conclusions.
First systematic comparison of agent architectures (tool-calling, MCP, code-generation, CLI) across heterogeneous environments and benchmarks.
PATRA framework improves LLM performance on time series question answering by capturing temporal patterns and balancing task complexity.