TIP: Token Importance in On-Policy Distillation
Token Importance analysis in on-policy knowledge distillation identifying which token positions provide most useful learning signals for student LLMs.
Token Importance analysis in on-policy knowledge distillation identifying which token positions provide most useful learning signals for student LLMs.
Contribution Weighted Group Relative Policy Optimization for training LLM-based search agents via reinforcement learning with improved credit assignment.
Best-Arm Identification framework for autonomous reasoning and planning with LLM agents, addressing systematic evaluation biases in search expansion.
Tabular foundation models for molecular property prediction using in-context learning without task-specific fine-tuning, improving drug discovery applications.
Knowledge-transfer network for multimodal sentiment analysis that reconstructs missing modalities during training and testing phases.
Research on logit suppression vulnerabilities in LLM safety alignment, introducing SSAG method to systematically identify and manipulate safety mechanisms in large language models.
Memory-aware techniques for efficient LLM training with automatic configuration in resource-constrained environments.
Automatic Dataset Construction methodology for large-scale data collection and curation with minimal human annotation.
SFTMix method improving LLM instruction tuning using mixup data augmentation without requiring data filtering.
Survey of plasticity loss in deep reinforcement learning, examining causes and solutions for adaptation failures.
Survey of distribution shift challenges and solutions in medical image analysis deep learning models.
Computational method for automating open code identification in qualitative analysis using generative AI tools.
Fine-tuning method to reduce LLM hallucinations through uncertainty calibration for improved trustworthiness.
Security research on context poisoning attacks against AI coding assistants, demonstrating vulnerabilities in automatic context gathering.
Multimodal retrieval-augmented system for chest X-ray report generation using key phrase extraction to reduce hallucinations.
Study evaluating 10 LLMs on code reasoning, distinguishing semantic understanding from lexical pattern matching in long context code understanding.
Study evaluating LLM alignment with human moral values in high-stakes decisions using kidney allocation scenarios.
ReGA: representation-guided abstraction technique to safeguard LLMs against jailbreak attacks and harmful content generation.
R3D2 uses diffusion models to insert photorealistic 3D assets into autonomous driving simulations from real-world data.
StableMTL repurposes latent diffusion models for multi-task learning from partially annotated synthetic datasets in zero-shot setting.
Security vulnerability where single users can persistently poison LLM knowledge through feedback upvoting/downvoting without external access.
Lizard linearization framework transforms Transformers into subquadratic architectures for efficient long-sequence LLM inference.
MetaLint: meta-learning framework for code linting that generalizes to unseen best practices using instruction-following LLMs.
Random Matrix Theory analysis of deformed weight matrices in DNNs, relating to neural network pruning techniques.
RAG framework with LLMs for diverse cross-cultural recipe adaptation considering dietary needs and cultural appropriateness.
VocabTailor dynamically selects vocabulary for small language models to reduce memory footprint on edge devices.
Knowledge distillation framework for text embeddings enabling lightweight student models aligned to larger teacher representations.
Efficient fine-tuning method focusing on bias terms of LLMs for parameter efficiency in low-data scenarios.
Investigation of how LayerNorm causes recency bias toward later tokens in Transformer decoders via causal self-attention interaction.
Analysis of fine-tuned LLM judges' temporal robustness, backward-compatibility, and generalization across question types.
Theoretical analysis of continual learning limitations using one-hidden-layer quadratic neural networks on sequential tasks.
Parallel test-time scaling for latent reasoning LLMs, comparing efficiency of continuous vector space reasoning versus explicit chain-of-thought.
Tensor compiler enabling advanced operator fusion for reduction operators including attention mechanisms on GPUs.
AttWarp method using attention-guided image warping to improve MLLMs' fine-grained spatial reasoning and small detail detection.
Training-free LLM-assisted watermarking for source code using semantic-preserving transformations to prevent unauthorized reuse.
Agentic training approach for multi-turn text-to-SQL with execution, verification, and refinement for coherent dialogue grounding.
Executable knowledge graphs as representations to improve LLM agent replicability of AI research through better code generation and background knowledge.
ML-assisted approach for solving quadratic unconstrained binary optimization problems, demonstrating practical advantages over classical methods.
Automated framework for discovering and evolving jailbreak attack strategies against LLMs with continuous learning capabilities.
Study of entropy collapse in reinforcement learning with verifiable rewards for improving LLM reasoning capabilities and avoiding premature convergence.
LoRA on the Go: Instance-level dynamic LoRA selection and merging for parameter-efficient fine-tuning across diverse domains without labeled data.
REFLEX: Reference-free evaluation metric for log summarization using LLM judgment to assess quality across relevance and informativeness dimensions.
SHRUG-FM: Reliability-aware geospatial foundation models detecting out-of-distribution failures in Earth observation.
VideoP2R: Process-aware reinforcement fine-tuning framework separating perception and reasoning for video language models.
FireScope: Large-scale dataset and chain-of-thought oracle for multimodal wildfire risk prediction from satellite imagery.
Systematic evaluation of off-policy training data impact on LLM behavior monitoring probes across eight behaviors.
Study investigating metadata types beyond URLs for efficient LLM pretraining, identifying document quality signals.
Matrix: Peer-to-peer multi-agent framework for coordinated synthetic data generation without centralized orchestrator.
Representational contrastive scoring method for efficient, generalizable jailbreak detection in Vision-Language Models.
ReASC: Reliability-aware adaptive self-consistency method reducing inference cost in LLM reasoning by reweighting response samples.