LLM-based method for schema-adaptive tabular representation learning to improve generalization across varying EHR schemas in clinical reasoning.
Framework for learned capability governance in autonomous AI agents, addressing overprovision of tool access across different task types.
Study on LLM value alignment via midtraining using synthetic documents; releases ANIMA evaluation dataset for animal compassion reasoning.
Single-layer Mamba architecture for time series classification with minimal design modifications.
Self-play framework with formal verification to improve LLM code reasoning in Haskell, includes 28k synthetic dataset.
Benchmark dataset for evaluating LLM performance on Indian financial regulatory documents, addressing non-Western text gap.
Privacy analysis: study of personality trait inference from ChatGPT chat history across 668 users.
ASP(Q) framework for inconsistency-tolerant querying of prioritized data with optimal repair semantics.
Study of memory tokens as computational scratchpad in Universal Transformers for adaptive recursive reasoning on combinatorial problems.
MetaErr: method for predicting error patterns and failure modes in deep neural networks.
PU-GKAN: Kolmogorov-Arnold network using partition-of-unity Gaussian basis functions with normalized activations.
Multi-agent constraint-guided framework for decompilation recovering executable source code from binaries.
Structured skill representation framework converting text-based skill descriptions into machine-usable scheduling, logic, and control structures for LLM agents.
G-Loss: graph-guided loss function incorporating semi-supervised label propagation for fine-tuning language models.
Neural cellular automaton for semantic parsing with structural generalization without hand-written compositional rules.
Mechanistic analysis of why LLM agents deviate from Nash equilibria in games and methods to reverse deviations.
Threat modeling of LLM-enabled robotic systems tracing attack propagation from prompts to physical actuation.
Analysis of why self-supervised encoders prefer normal distributions in joint-embedding predictive architectures.
CastFlow: multi-agent system using role-specialized LLM workflows for improved time series forecasting with iterative refinement.
Survey of LLMs in peer review automation: covers review generation, agent systems, RL methods, and future paradigms.
Attractor FCM: gradient descent-based fuzzy cognitive map using residual memory and backpropagation through time.
Black-box on-policy distillation for multimodal model alignment, addressing distributional drift in supervised fine-tuning.
SocialBias-Bench: benchmark studying social bias in LLM-generated code across 343 real-world tasks with mitigation strategies.
StateSMix: lossless compression algorithm using Mamba state space models and n-gram mixing without pre-trained weights.
eOptShrinkQ: KV cache compression for transformers using spectral denoising and quantization to reduce memory overhead.
OpsLLM: domain-specific LLM framework for software operations with knowledge-based QA and root cause analysis capabilities.
Analysis of systematic verification errors in RL with verifiable rewards for LLM reasoning improvement.
Agentic AI framework using MoE and LLMs for 6G network optimization and resource orchestration.
Comprehensive survey of rollout strategies in LLM reinforcement learning for reasoning and tool use.
Safety collapse in fine-tuned guard models for agentic AI pipelines through domain specialization.
Strategy-aware AI agent framework for autism intervention using LLMs trained on real clinical ABA data.
Self-supervised foundation model for cardiotocography analysis using multi-view SSL and clinical metadata.
Heterogeneous graph analysis and automated LLM-based interpretation for assessing urban bridge infrastructure importance.
Decentralized relay learning approach for sustainable large-scale machine learning training.
Reasoning foundation models for de novo protein design with explicit biochemical reasoning.
Training-free multi-agent framework for multimodal controversy detection in videos using audience perspectives.
Zero-shot interpretable multi-agent framework for detecting harmful content in memes without annotation.
Framework for analyzing intersectional bias in fetal ultrasound medical imaging tasks.
Comprehensive empirical study of multi-turn agentic RL for medical AI agents across clinical domains.
Study of pass-rate rewards in reinforcement learning for LLM code generation, addressing sparse reward problem.
Adversarial attack on Mixture-of-Experts LLMs exploiting routing mechanisms to bypass safety alignment.
Kernel Affine Hull Machines enabling efficient query-side semantic encoding without repeated neural inference for transformer retrieval systems.
ZeRO-Prefill optimization reducing distributed execution overhead in mixture-of-experts model serving for prefill-only discriminative tasks.
Bridge diffusion method with closed-form solutions for score functions and drift fields enabling analytical controlled path generation without neural networks.
ISAAC framework auditing causal reasoning in deep learning drug-target interaction models via intervention-based structural sensitivity probing.
Benchmark for evaluating LLM agents with tool use detecting reward hacking exploits through multi-step task evaluation with naturalistic shortcuts.
Diffusion-aided reward shaping approach for scheduling AIGC workloads across distributed data centers while minimizing energy costs.
AutoRAGTuner declarative framework automating RAG pipeline optimization through modular architecture and configuration-driven hyperparameter tuning.
Gradient transport analysis framework for understanding cascade efficiency and transport mechanisms during large language model pretraining.
Cross-lingual safeguard transfer framework improving multilingual safety alignment in LLMs through self-distillation from high-resource languages.