Rethinking LoRA for Privacy-Preserving Federated Learning in Large Models
Analysis of LoRA application in privacy-preserving federated learning for large vision and language models, revealing privacy-utility trade-offs.
Analysis of LoRA application in privacy-preserving federated learning for large vision and language models, revealing privacy-utility trade-offs.
DP-FedAdamW optimizer for federated learning of large models under differential privacy, addressing variance and convergence issues.
Study showing text-to-image diffusion models produce visually appealing but unreliable synthetic training data for computer vision tasks.
Red teaming framework for evaluating risks of LLMs in mental health support using AI psychotherapists and simulated patient agents.
Theoretical analysis showing equivalence between random network distillation, deep ensembles, and Bayesian inference for uncertainty quantification.
ReAttn method improving LLM-based re-ranking by re-weighting attention mechanisms to address concentration and shallow processing issues.
Safety framework (CORE) for robots reasoning about contextual safety requirements in open-world environments with varying conditions.
Federated learning framework addressing privacy, convergence, and Byzantine robustness in distributed model training without central server.
Semi-supervised method for detecting anomalies in dynamic graphs with limited labeled data, balancing unsupervised and supervised approaches.
AgenticSum: agentic inference-time framework using LLMs for clinical text summarization with multi-step verification to reduce hallucinations.
Robot navigation system that manipulates obstacles to find paths in cluttered environments using constraint-based planning.
World-model-driven diffusion policy framework with online adaptive learning for robotic manipulation under dynamic conditions.
Formal analysis of information flow and attack surface in AI agent systems, showing how malicious prompt injection compromises multi-step conversations.
Study benchmarking LLM comprehension across multilingual datasets, showing significant performance gaps in low-resource non-Western languages.
Zero-shot vision-language framework estimating urban heat demand from satellite imagery using pretrained LVLM features.
Policy gradient method for multi-agent reinforcement learning reducing cross-agent noise variance for scalable cooperative learning.
BarrierSteer framework for LLM safety using barrier functions and steering, addressing adversarial attacks and unsafe content generation.
Benchmark for machine unlearning methods on Vision Transformers, extending prior CNN-focused unlearning research to transformer architectures.
NovaPlan hierarchical framework combines vision-language models with video planning and geometric robot control for long-horizon manipulation tasks.
NanoKnow analyzes how LLMs encode knowledge by studying nanochat models with fully open, transparent pre-training data.
Selective Chain-of-Thought improves medical QA efficiency by predicting whether questions require reasoning before generating rationales.
AdaEvolve uses LLMs as semantic mutation operators in evolutionary optimization with adaptive scheduling for automated program generation.
KNIGHT framework generates multiple-choice question datasets from knowledge graphs and external sources using LLMs for RAG system evaluation.
AgentOptics framework enables agentic AI for autonomous optical system control using Model Context Protocol with 64 MCP tools and 410-task benchmark.
Behavior Learning framework learns interpretable hierarchical optimization structures from data with identifiability guarantees.
Benchmark suite evaluating video model reasoning capabilities on spatiotemporal understanding including continuity, interaction, and causality.
WorldGUI benchmark evaluates GUI agents on desktop automation tasks from arbitrary starting points with partial configurations and state variability.
Ambig-SWE studies LLM agent capabilities to handle underspecified instructions by asking clarifying questions in interactive software engineering tasks.
V-Droid is a mobile GUI automation agent using LLMs as verifiers to evaluate candidate actions before execution, with discretized action space framework.
Novel meta-continual learning approach for neural fields using modular architecture and optimization-based meta-learning to address catastrophic forgetting.
MOFh6 uses LLMs to extract and standardize metal-organic framework synthesis conditions from raw articles and crystal codes into structured tables.
Theoretical analysis of top-k decoding method for LLM sampling, validating sparsity assumptions in next-token distributions.
Relational first-order representation (Foreplan) for efficient forward planning in factored MDPs with concurrent actions and objects.
SOP-Bench: 2000+ task benchmark for evaluating LLM agents on complex multi-step industrial procedures across 12 business domains.
Decision-theoretic framework for evaluating model explanations based on actual decision task performance improvement.
AI framework using multimodal social media data to analyze tourist perception and preferences in historic urban areas.
Theoretical analysis comparing linear programming vs. dynamic programming approaches for solving Markov decision problems.
Neuro-symbolic geometry theorem prover combining neural diagram understanding with visual language models and symbolic verification.
ImitSAT: Branching policy for SAT solvers using imitation learning from expert traces instead of reinforcement learning.
DIVER: Diversity-incentivized exploration method for RL with verifiable rewards to improve LLM reasoning sample efficiency.
Multimodal embedding model with visual-interactive capabilities for region-of-interest specification in vision-language tasks.
Method leveraging offline trajectories to improve LLM web agent adaptation without costly online interactions or fine-tuning.
Framework conceptualizing AI agents as stochastic dynamical systems for learning reasoning via transductive inference on new tasks.
Analysis of design choices in RL-based deep research agents augmented with external tools for complex web-based question answering.
Comprehensive review of large-scale AI models applied to neuroscience domains including brain-computer interfaces and neural decoding.
Clinical-grounded dataset and multi-agent dialogue framework for psychiatric comorbidity using synthetic EMRs and LLM agents.
First visual backdoor attack framework (BEAT) targeting vision-language model embodied agents via contrastive trigger learning.
Benchmark comparing causal vs. correlation-based approaches for predictive maintenance in manufacturing with extreme cost asymmetry.
Evaluation framework testing LLM reasoning robustness under rule-based perturbations via stress tests on logic consistency.
Safety benchmark evaluating outcome-driven constraint violations in autonomous AI agents deployed in high-stakes environments.