Structured Agent Distillation for Large Language Model
Structured Agent Distillation compresses LLM-based agents into smaller student models while preserving reasoning and action consistency.
Structured Agent Distillation compresses LLM-based agents into smaller student models while preserving reasoning and action consistency.
Framework for verifying correctness of math questions used in LLM training. Focuses on QA data quality beyond answer correctness.
AudioTrust benchmark evaluating trustworthiness of audio LLMs. Reveals vulnerabilities from non-semantic acoustic cues like timbre and accent.
Steganographic jailbreak attacks on LLMs balancing semantic and linguistic stealth. Bypasses safety mechanisms through hidden malicious intent.
ReasonMap benchmark for evaluating multimodal LLM visual reasoning on transit maps. Tests math and logic capabilities on 1,008 questions.
Data-driven survey of 14,648 papers on LLM limitations from 2022-2025. Systematically categorizes known weaknesses and failure modes.
Investigates LLM limitations in theoretical physics. Identifies gaps in physical intuition and constraint satisfaction beyond prompting improvements.
Study measuring how well LLMs comprehend user intent beyond surface-level text matching. Analyzes gap between token prediction and actual user goals.
Refine-POI applies reinforcement fine-tuning to LLMs for point-of-interest recommendation with improved semantic ID indexing and topology awareness.
Research on adapter parameters and task merging for efficient multi-task learning in on-device LLMs, enabling multiple tasks via parameter merging.
TURA proposes a tool-augmented retrieval agent for conversational AI search that handles real-time data and structured queries beyond traditional RAG limitations.
Once4All uses LLM-synthesized test generators guided by skeleton templates to fuzz SMT solvers and uncover correctness bugs.
Fast Image-to-Neural Surface constructs implicit distance representations from single images for robotics obstacle avoidance and path planning.
DiDi-Instruct distills fast student models from diffusion LLMs for ultra-fast language generation matching teacher performance.
TRACE uses AI for semi-automated assessment of individual contributions in collaborative computer science group projects.
XGrasp detects robotic grasps that generalize across multiple gripper types without retraining using gripper-aware architecture.
DriveCritic framework uses vision-language models to provide context-aware evaluation of autonomous driving planners aligned with human judgment.
Vision-language model approach for 3D spatial reasoning from limited views using geometric imagination grounding.
Unifying framework explaining in-context learning and activation steering through belief dynamics, treating both as instances of broader control mechanism.
Study evaluates non-functional quality characteristics of LLM-generated code using ISO/IEC 25010 model across functional correctness, maintainability, and security.
DeepSport is an end-to-end trained multimodal LLM for multi-sport video understanding using agentic reinforcement learning for iterative reasoning.
ConCISE is a reference-free evaluation metric for measuring conciseness of LLM-generated responses to reduce verbosity and token costs.
MedEyes applies vision-language models with reinforcement learning for medical diagnosis via dynamic visual focusing and iterative clinical reasoning.
Knowledge distillation method using structured chain-of-thought to improve text-to-SQL performance in small language models while reducing cost and security risks.
KnowVal combines visual-language reasoning, driving knowledge, and value alignment for autonomous driving using knowledge-augmented approaches beyond data-driven learning.
Dataset (judgeWEL) for named entity recognition in Luxembourgish using LLMs to verify automatically labeled data for underrepresented languages.
Benchmark and framework (Grand-SMOT) for semantic multi-object tracking using multimodal LLMs to handle complex relational queries.
arXiv: LMMRec framework using LLMs to model user motivations in multimodal recommendation systems from heterogeneous information sources.
arXiv: LatentChem decouples chemical reasoning from natural language by using latent representations instead of chain-of-thought prompting.
arXiv: Hybrid-policy reinforcement learning approach for enhancing multi-modal LLM reasoning with controlled exploration strategies.
arXiv: Evaluating small language models for zero-shot and one-shot role classification in robot leader-follower interaction tasks.
arXiv: FlashOptim memory-efficient optimizers for mixed-precision neural network training reducing parameter storage overhead.
arXiv: ProtoDCS framework for robust test-time adaptation of Vision-Language Models under distribution shift and open-set scenarios.
arXiv: Using LLMs with structured prompts to generate auxiliary lemmas for constraint solving with inductive definitions.
Study of parental moderation preferences for children's GenAI chatbot interactions using LLM-generated probe scenarios.
BLOCK: open-source two-stage MLLM pipeline generating Minecraft character skins from text descriptions via 3D preview synthesis.
Multi-dimensional evaluation of LLM safety benchmarks assessing their academic influence and code repository quality.
Analysis of performative chain-of-thought in reasoning LLMs showing discrepancy between generated explanations and internal model beliefs.
Design study on LLM-assisted tools for systematic literature reviews to reduce cognitive load and enable iterative exploration.
Investigation of tokenizer pretraining impact on physics foundation models for emulating complex multiphysics simulations.
Expert perspectives on integrating foundation models and AI agents into clinical computational pathology with translational readiness assessment.
Framework using structured perturbations to evaluate LLM performance on high-stakes grant proposal reviews across quality dimensions.
EXPLORE-Bench benchmark evaluating multimodal LLMs' ability to reason about long-horizon physical consequences from egocentric viewpoints.
AraModernBERT adapter for Arabic NLP with transtokenized embeddings and 8K token context window support.
Dataset creation using Wikidata to detect cultural biases in LLMs across Latin American languages and contexts.
Semantic embedding injection in neural transducers for low-latency streaming automatic speech recognition.
Epistemic Support-Point Filter for recursive estimation using maximum entropy and falsification principles.
Proposes harmonic loss as alternative to cross-entropy for training deep neural networks with improved interpretability.
Method to provably compute adversarial examples for black-box neural networks using Contract And Conquer approach.
Human-inspired reasoning approach for robust speech deepfake detection with improved generalization to unseen audio domains and interpretability.