Agentic AI-assisted coding offers a unique opportunity to instill epistemic grounding during software development
GROUNDING.md framework for epistemic grounding in agentic AI-assisted software development using field-scoped documentation.
GROUNDING.md framework for epistemic grounding in agentic AI-assisted software development using field-scoped documentation.
StructMem: structured memory system for LLM agents enabling long-horizon behavior through relational event modeling.
Analysis of cultural and regional biases in LLMs, revealing preferences for Japanese culture and Western-centric viewpoints.
Transient Turn Injection attack technique exploiting stateless moderation in multi-turn LLM interactions using automated agent-based adversarial testing.
GiVA parameter-efficient fine-tuning method using gradient-informed bases to improve vector-based adaptation over LoRA.
Analysis of hallucinations in large vision-language models showing how prompts can override visual grounding in outputs.
Comprehensive survey of evaluation methods for LLM-based agents covering planning, tool use, benchmarks, and interaction with dynamic environments.
C-SHAP method for explainability of time series models in high-stakes domains like healthcare and industry.
Bandit learning approach for adaptive test-time compute allocation to LLM queries based on estimated difficulty rather than uniform scaling.
KompeteAI autonomous multi-agent system for ML pipeline generation addressing exploration and execution bottlenecks in LLM-based AutoML.
PosterForest multi-agent system for hierarchical scientific poster generation using tree-structured planning for content and layout organization.
AIRL framework learning reasoning reward models from expert demonstrations for LLM improvement, bridging supervised fine-tuning and outcome-based RL.
Speculative Actions framework for accelerating AI agents through parallel action generation, reducing sequential API latency in interactive environments.
OpenEstimate benchmark evaluating LLMs on reasoning under uncertainty using real-world data from healthcare, finance, and knowledge work domains.
ReProbe method for efficient test-time scaling of LLM reasoning by probing internal states instead of expensive process reward models.
Philosophical analysis arguing static value alignment approaches are insufficient for AI systems under capability scaling and distributional shift.
arXiv benchmark (AgencyBench) evaluating LLM agents on long-horizon real-world tasks with 1M-token context and automated evaluation.
arXiv framework (AgentDoG) for diagnosing and mitigating safety/security risks in autonomous AI agents with unified taxonomy.
arXiv paper on LLM-based synthesis of game design patterns and executable gameplay code using design knowledge representations.
arXiv benchmark comparing LLM agent communication protocols for task orchestration, tool integration, and multi-agent delegation.
arXiv paper analyzing mathematical relationships between Bayesian networks and structural causal models with uncertainty.
arXiv survey of abductive reasoning in LLMs, covering inference of plausible explanations across diverse task settings.
arXiv benchmark (DRBENCHER) for evaluating deep research agents combining web browsing with multi-step computation tasks.
arXiv paper on QuarkMedSearch, an agentic AI system for medical domain search with multi-hop data construction and evaluation.
arXiv catalog of 195 AI safety benchmarks with structured metadata revealing fragmentation in LLM safety measurement practices.
arXiv benchmark (HWE-Bench) for evaluating LLM agents on real-world hardware bug repair at repository scale with 417 task instances.
arXiv benchmark (ReactBench) evaluating multimodal LLMs on topological reasoning over complex chemical reaction diagrams.
arXiv paper proposing visualization methods for examining distributions of LLM outputs instead of single samples.
FSFM proposes biologically-inspired selective forgetting mechanism for LLM agents based on hippocampal consolidation and forgetting curves for memory management.
Framework leveraging foundation model priors to enable efficient reinforcement learning for robotic manipulation tasks with minimal environment interaction.
Model-Agnostic Self-Decompression prevents catastrophic forgetting in LLMs during fine-tuning and addresses performance degradation in multimodal models.
FedCoLLM proposes federated parameter-efficient framework for co-tuning large and small language models with mutual enhancement.
Analytical FFN-to-MoE Restructuring converts dense LLM feed-forward networks to Mixture-of-Experts without extensive retraining, reducing inference costs.
SafeMERGE preserves safety alignment during LLM fine-tuning via selective layer-wise model merging, maintaining task utility without custom algorithms.
LogiBreak presents black-box jailbreak method using logical expression translation to bypass LLM safety mechanisms, analyzing distributional discrepancies in alignment.
Secure LLM Fine-Tuning via Safety-Aware Probing addresses how fine-tuning compromises LLM safety alignment and proposes detection methods.
Counterfactual Segmentation Reasoning method diagnoses and mitigates pixel-grounding hallucinations in Vision-Language Models for segmentation tasks.
mGRADE proposes lightweight sequence modeling architecture balancing local and global context under memory constraints for edge devices.
DiffuMeta combines diffusion transformers with algebraic language models for inverse design of 3D metamaterials, addressing computational complexity in material discovery.
Comprehensive survey of differential privacy theory, applications, and gap between formal guarantees and user expectations in ML and data security contexts.
HyperAdapt introduces a parameter-efficient fine-tuning method that reduces trainable parameters while adapting foundation models to specialized tasks with lower memory and compute requirements.
InfiniPipe proposes elastic pipeline parallelism to reduce communication overhead in long-context LLM training, addressing memory and efficiency tradeoffs between sequence and token-level parallelism.
Defines effective context window for LLMs and creates standardized testing method across models and problem types.
Reinforcement fine-tuning for geospatial referring expression understanding with limited labeled data.
Studies cross-modal reasoning in MLLMs and when modality interactions help or harm performance.
Chess-based testbed for evaluating strategic reasoning and rule adherence in LLMs.
Framework for active learning with unaligned multimodal data to reduce annotation costs.
First federated learning benchmark for surgical video analysis without sharing patient data.
Adaptive patch sizing for Vision Transformers reduces tokens via content-aware allocation.
Curriculum RL framework addressing performance degradation in multi-turn conversations with verifiable rewards.