FedPSA: Modeling Behavioral Staleness in Asynchronous Federated Learning
Addresses staleness challenges in asynchronous federated learning with behavioral modeling.
Addresses staleness challenges in asynchronous federated learning with behavioral modeling.
Fine-tunes LLMs to generate corrective transmission switching actions for power grid operations.
Integrates neural networks with symbolic reasoning for knowledge graph question answering using active exploration.
RL architecture inspired by cerebellar circuits to improve sample efficiency, noise robustness and generalization.
Framework addressing modality discrepancies between synthetic and real images for training machine learning models.
Evaluates LLM performance on slang in Australian and Indian English across 7 models using new datasets.
Framework for orchestration-free customer service automation using Task-Oriented Flowcharts to guide end-to-end LLM-based workflows without manual intervention.
Study of action tokenization design for Vision-Language-Action models, examining what makes effective action tokenizers for VLA optimization.
Mathematical analysis of representational similarity in discriminative models including autoregressive language models using logit distance bounds.
Framework combining deep generative models with quantum annealing for molecular design beyond training data distribution.
SecCodeBench-V2 benchmark for evaluating LLM code generation security across 98 scenarios and 22 CWE categories in 5 languages.
Memory framework for embodied AI agents using episodic-semantic separation to improve long-horizon question answering and exploration.
Training-free dynamic fusion framework for combining multiple LoRAs to generate subjects and styles in image generation.
VLM-DEWM cognitive architecture for vision-language planning in manufacturing with persistent state tracking and interpretable reasoning.
Dynamic workflow learning system for Text-to-SQL that adapts execution paths at inference time instead of static pipelines.
STAPO method stabilizing RL fine-tuning for LLMs by filtering spurious tokens to prevent late-stage training collapse.
Security analysis of self-evolving LLM agents where persistent memory can be exploited for covert payload injection attacks.
Bayesian optimization pipeline for selecting and tuning deep learning models for 3D biomedical image segmentation and classification.
Framework studying how neural networks represent latent geometry in dynamical systems through representational alignment analysis.
Content-based framework for cybersecurity refusal decisions in LLMs and agents, addressing consistency and brittleness issues.
Open-source simulator (LSMART) for evaluating multi-agent path finding algorithms in fleet management with automated guided vehicles.
LLM-based navigation agent for vision-and-language tasks that learns to retrieve navigable candidates efficiently in unseen environments.
Benchmark for evaluating multi-turn chart editing in multimodal language models with iterative refinement.
Framework addressing trade-offs between generative and understanding capabilities in multimodal models via Reason-Reflect-Refine algorithm.
Analysis showing fine-tuning causes alignment collapse in language models and proposing geometric understanding of safety degradation.
Decision quality evaluation framework for assessing content moderation by both human agents and LLMs at scale.
Encoder-only adaptation of Avey attention-free architecture for efficient NLP under compute constraints.
Scalable non-destructive LLM editing method using second-order constraints to preserve capabilities while modifying target behaviors.
LLM-based in-context learning for optimizing public safety UAV navigation and control during emergency response.
AI agent framework integrating LLMs with Lean formal proof assistant for automated theorem proving with auxiliary lemma generation.
Comprehensive safety evaluation framework for real-world AI agents across diverse task domains and tool interactions.
LLM-based in-context learning approach for UAV resource allocation in wildfire monitoring systems.
Shapley-based explanation framework for understanding cross-modal interactions in multimodal AI models.
Framework for LLM agents to learn when to plan via reinforcement learning, improving efficiency and performance on long-horizon tasks.
Theoretical analysis establishing accuracy upper bounds for LLM single-pass reasoning in multi-hop question answering, formalizing capacity overflow bottlenecks.
Identifies and mitigates priming vulnerability in diffusion language models exploited by jailbreak attacks during iterative denoising inference.
Generalized parallel LLM inference scaling uses interdependent generations to share information across parallel responses, improving quality and efficiency.
SR-Scientist elevates LLMs from equation proposers to autonomous AI scientists for scientific discovery, writing and executing code for symbolic regression.
Studies expressivity of structured argumentation frameworks with uncertainty in rules and premises, advancing formal argumentation theory.
Generative semantic workspaces enhance RAG with episodic memory for long-context reasoning, addressing context window limitations and sequence length degradation.
Multi-agent AI framework reconstructs pre-crash vehicle scenarios from fragmented collision data using collaborative reconstruction and reasoning stages.
Chain of Summaries generates information-dense summaries of web content via iterative questioning, enabling LLMs to process diverse content formats within context limits.
Aeon proposes neuro-symbolic memory management for long-horizon LLM agents, using hierarchical structures to overcome context window and attention limitations.
ErrorMap charts failure sources in LLM benchmarks, disentangling causes like formatting errors and calculation mistakes from reasoning deficiencies.
PolySHAP improves Shapley value computation for explainable AI using polynomial regression, reducing computational cost of KernelSHAP approximation.
ScholarGym benchmarks LLM capabilities in iterative research workflows, evaluating information-gathering stages of deep research systems beyond end-to-end metrics.
FlowSteer: end-to-end reinforcement learning framework for agentic workflow orchestration. Learns policies for multi-step task automation with minimal manual specification.
MARS: AI agent framework for automated machine learning research. Budget-aware planning, modular script generation, and reflective search for computationally expensive experiments.
LQA: lightweight framework for deploying vision-language models on edge devices. Combines quantization with gradient-free test-time adaptation for distribution shift robustness.
stable-worldmodel-v1: reproducible open implementation of world models for agent reasoning and planning. Standardizes evaluation and reduces publication-specific code fragmentation.