Multi-agent reasoning framework for automated software system performance optimization beyond local code transformations, reasoning about whole-system interactions.
AdapterTune adds zero-initialized low-rank adapters to frozen Vision Transformers for stable transfer learning with principled capacity guidance.
Privacy-preserving machine translation at inference stage with new benchmark dataset for evaluating local translation without cloud servers.
POLCA framework uses LLMs as optimizers guided by rewards and feedback to automate optimization of prompts and multi-turn agent systems, formalizing it as stochastic generative optimization.
Hybrid-order split federated learning combining zeroth-order optimization with standard backprop for memory-efficient fine-tuning.
Privacy-preserving RAG service supporting arbitrary top-k retrieval for LLM-based systems with secure document retrieval.
Framework for lifelong agents requiring epistemic control to select appropriate reasoning frameworks and prevent decision chain failures.
Probabilistic certification method for verifying behavioral fidelity in compressed deep neural networks.
Model-agnostic unlearning framework using ratio-aware layer editing for vision transformers and diffusion models.
Inference-time feature projection balancing safety and utility tradeoffs in large vision-language models.
Causal analysis of residual stream hyper-connections in multi-stream transformer architectures exploring mechanistic interpretability.
Continual learning framework for toxicity detection adapting to evolving evasive perturbations in online content.
LLM-based decompilation tool translating pseudocode to compilable executable code with runtime correctness verification.
Architecture-agnostic defense mechanism against heterogeneous generative threats including diffusion models and GANs.
Sample-efficient hypergradient estimation for decentralized bi-level reinforcement learning with leader-follower agent dynamics.
LLM-based automated essay scoring using decision-level ordinal modeling for multimodal inputs with trait-specific visual relevance.
Signal Detection Theory analysis of LLM calibration metrics, decomposing sensitivity and bias components beyond ECE.
Post-hoc explanation method using informative perturbation selection for uncertainty-aware model interpretability.
Lightweight routing mechanism for transformer attention heads with mechanistic interpretability analysis of computational pathways.
Framework for question-aware keyframe selection in video question answering using synthetic supervision.
Multi-agent simulation framework generating synthetic corporate corpora with verifiable ground truth for RAG pipeline evaluation.
Framework and architectural patterns for describing agentic AI systems with multi-agent collaboration and artifact exchange.
Replication study on AI-generated text detection using multilingual models and SHAP-based explainability analysis.
Vision-Language-Action model using state space models for efficient language-guided robotic manipulation tasks.
Latent-space reasoning approach for LLMs that reduces inference cost by computing reasoning steps implicitly rather than generating verbose traces.
Analysis of error sources in global feature effect estimation methods like PD and ALE plots for model interpretation.
Open-source biomedical knowledge graphs with federation and AI agent access for cross-referencing siloed databases.
Research on improving LLMs' code generation using private libraries through better knowledge integration beyond API documentation retrieval.
HindSight evaluation framework measuring AI-generated research idea quality by matching against future publications and citation impact.
Hybrid approach combining Iterative Learning Control with deep reinforcement learning for safe and convergent batch process control.
Token Coherence framework applying MESI cache protocols to reduce synchronization overhead in multi-agent LLM orchestration systems.
CATFormer combines continual learning with spiking transformers using dynamic thresholds to mitigate catastrophic forgetting.
In-context symbolic regression for extracting interpretable analytical expressions from Kolmogorov-Arnold Networks in scientific ML.
RESTA defense extended to vision-language models for robustness against multi-modal jailbreaking attacks in trustworthy agentic AI.
Code-centric learning approach for LLM-based ICD medical coding improving generalization to unseen codes with better interpretability.
Scalable simulation-based model inference framework with test-time complexity control for selecting among large families of forward models.
CCTU benchmark evaluating LLM tool use under complex constraints, testing function calling, instruction following, and self-refinement.
SKILLS benchmark framework evaluating LLM agents on 37 telecom operations workflows with real API interfaces, testing structured knowledge injection.
Hybrid XAI framework combining counterfactual explanations and feature attribution for neural network interpretability in healthcare/finance.
Analysis of beam search in LLMs showing wider beams can hurt output quality due to overestimation bias, grounded in Extreme Value Theory.
Spatial reasoning agent decoupling perception from reasoning in visual language models for improved metric and geometric scene understanding.
Bi-level optimization approach using Stackelberg game theory for coupled morphology-control co-design in embodied agents.
Adversarial patch framework for evasion and impersonation attacks against facial re-identification systems across non-overlapping cameras.
Safety defense mechanism for LLMs monitoring intermediate reasoning steps in chain-of-thought to prevent jailbreak attacks.
Benchmark evaluating marginal utility of agent skills for LLM-based software engineering agents on real GitHub issues and requirements.
Empirical study of 16 LLMs examining internal mechanisms for table understanding across attention dynamics, layer depth, and expert activation.
Comprehensive safety evaluation and monitoring framework for LLM-based multi-agent systems addressing novel risks beyond single agents.
Framework for improving robustness of quantized DNNs through three-stage fine-tuning addressing both fault and attack resilience.
Analysis of safety vulnerabilities in test-time training methods for LLMs, examining susceptibility to prompt injection and adversarial attacks.
Vision-language critic model leveraging pre-trained VLAs for multi-agent reinforcement learning value estimation with improved generalization.