StateLinFormer: Stateful Training Enhancing Long-term Memory in Navigation
StateLinFormer: linear-attention navigation model with persistent memory for long-term navigation tasks, combining flexibility with efficiency.
StateLinFormer: linear-attention navigation model with persistent memory for long-term navigation tasks, combining flexibility with efficiency.
Dual-Criterion Curriculum Learning proposes a meta-learning approach using dual criteria for difficulty assessment in temporal data training.
APreQEL proposes adaptive mixed precision quantization to reduce memory and computational costs of LLMs for edge device deployment while maintaining performance.
Time-LLM model for predicting wafer-level spatial etch depth distributions in plasma etching process monitoring.
LLMORPH automated testing tool for LLMs using metamorphic testing to detect NLP task failures without human-labeled oracles.
LLMLOOP framework automating iterative refinement of LLM-generated code and test cases through automated feedback loops.
Theory of LLM information susceptibility analyzing fundamental limits of LLM-mediated optimization in agentic systems.
Swiss-Bench SBP-002: trilingual benchmark of 395 expert-crafted regulatory compliance tasks across FINMA, Legal-CH, and EFK domains.
Probing study revealing how LLMs internally represent different ethical frameworks with asymmetric transfer patterns across model sizes.
Training-free out-of-distribution detection using multi-layer prototype fusion approach for robust deep learning deployment.
Privacy-preserving LLM system for disambiguating clinical acronyms in healthcare without transmitting data to external servers.
Measurement methodology for identifying assessment items where LLMs perform differently than humans using theory-grounded evaluation.
Analysis of early-exit decoding in modern LLMs showing reduced efficiency gains due to improved architectures with lower layer redundancy.
Study of filtered vector search algorithms in PostgreSQL for semantic search and GenAI applications, evaluating real-world database performance.
Self-paced curriculum learning for RL using closed-form Gaussian updates to improve efficiency in high-dimensional contexts.
Intent-Based Networking using AI to translate high-level natural language intents into network policies with automated compliance assurance.
Bayesian latent transport framework for domain-adaptive foundation models addressing distribution mismatch and uncertainty propagation in limited-supervision scenarios.
Cognitive Firewall: hybrid edge-cloud architecture for securing browser-based LLM agents against indirect prompt injection attacks using split-compute security checks.
LLM-informed model-based planning for object search using LLM likelihood estimates and environment costs.
Neural Regression Collapse phenomenon across network layers showing feature sparsity and low rank structures.
AgentPex: Framework for detecting procedural failures in agentic traces including workflow routing and tool usage violations.
Method for finding representations in language models via adversarial perturbation without implausible constraints.
Benchmark (PoliticsBench) measuring political bias in eight LLMs using multi-turn roleplay evaluation.
Research on activation function curvature role in adversarial robustness using Recursive Curvature-Tunable Activation Family.
Discussion of user experience design for generative AI in education emphasizing human-AI epistemic partnership.
Investigation of vision-language model robustness under distribution shifts using visual deductive reasoning tasks.
HDPO method augmenting RL with privileged self-distillation for LLM mathematical reasoning on unsolvable cliff prompts.
Luna: C++ implementation of alpha-CROWN bound propagation for neural network formal verification.
Multi-agent robotic platform using AI agents for adaptive chemical laboratory automation handling diverse experimental tasks.
Self-distillation method for multi-token prediction in LLMs to improve inference efficiency and MTP head acceptance rates.
Multimodal deception detection system using schema-driven approach with multicultural datasets and explainable reasoning.
MVH-Bench dataset and analysis of multi-view hallucination in vision-language models processing diverse viewpoint images.
LLM-enabled framework for automated threat hunting using Splunk SOC logs to assist security analysts with APT detection.
Systematic study of reasoning LLM inference costs revealing pricing reversal phenomenon where cheaper models cost more across 9 diverse tasks.
Physics-guided text-to-motion framework for humanoid control using rectified flow and safety gating to prevent kinematic hallucinations.
Ensemble of specialized LLMs architecture for adaptive tutoring that separates pedagogical decision-making from response generation.
Analysis of challenges in iterative generative optimization with LLMs for self-improving agents, identifying hidden design choices limiting adoption.
Fine-tuned 8B model for text-to-SQL at scale, reducing API costs and latency for production deployment in conversational applications.
Training framework addressing contextual exposure bias in speech-LLMs using teacher error knowledge and contrastive learning.
Method to reduce object hallucinations in LVLMs by rectifying attention imbalance across and within vision-language modalities.
Safety analysis of MLLMs for image generation, identifying semantic understanding capabilities that may introduce new risks compared to diffusion models.
Multi-task robotic manipulation framework using knowledge graphs and dynamic relation mechanisms for vision-grounded policy learning.
Dual-guidance RL framework for LLMs that combines external execution feedback with internal experience for improved reasoning task learning.
Comparative study of dual-form attention networks for multi-modal satellite time series analysis in land monitoring applications.
Analysis of response homogenization in RLHF-aligned LLMs and its impact on uncertainty estimation methods, identifying alignment-robustness tradeoffs.
Multilingual multi-turn medical dialogue dataset for training conversational AI systems in healthcare with improved realism and accessibility.
Scaling RL for LLM code generation using synthetic data pipelines and curriculum learning, addressing data diversity over volume.
Security analysis of Model Context Protocol (MCP) tool-augmented LLM agents, demonstrating stealthy injection attacks on tool responses.
Knowledge distillation method using dual-modality (vision + text/CLIP) teacher models to improve student model efficiency and quality.
Privacy analysis of time series imputation models, demonstrating membership inference and attribute leakage vulnerabilities in black-box settings.