Unified benchmark (PepSpecBench) for peptide tandem mass spectrometry prediction addressing evaluation challenges in deep learning for computational proteomics.
Phone2Act: Hardware-agnostic teleoperation framework using smartphone and ARCore for low-cost, scalable vision-language-action data collection across robot platforms.
TRAP attack targeting world-model planning agents through tail-aware ranking manipulation to compromise long-horizon planning and generalist agent behavior.
Trojan Hippo attack exploiting LLM agent memory systems for data exfiltration through dormant payloads planted via single untrusted tool interactions.
First large-scale reproducible benchmark for machine learning on Raman spectroscopy, standardizing evaluation across fragmented datasets and spectral models.
AI agent system for automating legal judgment drafting using agentic information retrieval and rubric-guided generation to improve evidence recall and reasoning.
Training-free LLM approach with prompt engineering for conventional commit classification, eliminating need for large labeled datasets and ML models.
Multimodal dataset addressing ambiguity resolution in machine translation by combining visual grounding with text, identifying data quality issues in prior benchmarks.
Low-cost modular robotic platform (VILAS) for vision-language-action policy learning and deployment with integrated arm, gripper, and dual-camera perception.
Multi-variant reliability audit of 15 open-weight language models showing single-prompt accuracy misses important failure modes in calibration and robustness.
Framework for standardizing AI evaluation through randomized controlled trials, adopting validity principles from established experimental disciplines.
Coopetition-Gym v1 benchmark platform for mixed-motive multi-agent RL with twenty environments covering interdependence, trust, collective action, and sequential dynamics.
Pair2Score transfers pairwise LLM comparisons into absolute essay scoring predictions using parameter-efficient LLaMA adaptation in two-stage framework.
Automatic joint pruning and quantization method for 3D Gaussian splatting compression to reduce model size from gigabytes to mobile-friendly formats.
STABLEVAL framework for disagreement-aware evaluation of AI systems that models annotator reliability and item ambiguity instead of using majority vote.
Theoretical analysis of boundary mass and soft-to-hard routing transitions in mixture-of-experts models under population-level squared-loss regression.
Characterizes sample complexity of offline multi-armed bandits with KL regularization, providing sharp bounds for KL-regularized performance metrics.
Survey consolidating fragmented literature on reusing trained models in deep RL through transfer, distillation, ensemble, and federated training methods.
Agentic system using LLM critic-guided reflexion to automatically maintain software documentation consistency as codebases evolve, reducing API misuse.
Improves Integrated Gradients feature attribution for neural networks by using manifold-aligned guidance to reduce unreliable explanations in noisy gradient regions.
Identifies security vulnerabilities in Bring-Your-Own-Key LLM agent architectures where malicious relays can tamper with aligned LLM responses post-generation.
Research on training data requirements for long-horizon software engineering AI agents, arguing for multi-engineer collaborative data instead of larger GitHub scrapes or solo trajectories.
Unified ablation study of LLM privacy attacks (MIA, AIA, DEA, backdoor) across system factors, revealing behavior under common deployment conditions.
System-level analysis of LLMs for multilingual orthopedic diagnosis in low-resource settings, evaluating reliability, calibration and safety for clinical decision support.
Efficient LiDAR-based place recognition using Bird's Eye View representations for edge deployment in autonomous navigation systems.
HELIX method for time series imputation using learnable feature identities as persistent anchors to maintain consistent cross-dimensional representations.
Analysis of slot collapse failure mode in attention-based compositional inference models and residual evidence modeling approach.
Framework for LLM-enabled social agents with grounding in roles, norms, intentions and contextual constraints for meaningful social interaction.
Autonomous vulnerability management system using LLM agents for bare-metal industrial OT networks, requiring protocol-level reasoning without shell access.
Study of structured output reliability gap in small LMs (7-9B), evaluating JSON format compliance and mathematical correctness under different prompting strategies.
Privacy-preserving ML workflow combining anonymization with personalized differential privacy budgets in federated learning architectures.
Analysis of in-context learning limitations in vision-language models and methods to close inductive gaps through inductive-deductive reasoning.
Vision for causal inference methods in software engineering to support AI-driven decision-making and LLM-based agents with interventional and counterfactual reasoning.
Structured spec-driven engineering (SSDE) approach using artifacts to guide LLM code generation at repository level, improving on function-level generation.
Statistical framework (CAFE) for studying multi-agent LLM systems under semantic stress to support antifragile learning rather than just robustness.
ML research on reinforcement learning with verifiable rewards, combining weighted SFT with policy optimization via reference-sampled Boltzmann projection.
Multi-agent retrieval-augmented AI framework for exploring and interpreting high-energy physics literature and experimental data.
Study on preference poisoning attacks against offline RLHF methods like Direct Preference Optimization.
Research investigating transfer learning from sleep biosignal pretraining to non-sleep medical tasks.
Preprocessing-driven approach using temporal convolutional networks for remaining useful life prediction in aero-engines.
Systematic empirical study comparing five retrieval strategies for biomedical retrieval-augmented generation systems.
Framework integrating vision-language models with indoor mobile robots via semantic reasoning and cross-robot adaptive memory.
Method for synthesizing neural barrier certificates to verify safety of dynamical systems using set-based training.
Framework for 3D indoor scene generation using zone-graph paradigm instead of object-centric synthesis.
Reinforcement learning approach for chemotherapy dose optimization under partial observability using memory-augmented policies.
Research on open-set panoptic segmentation using hierarchy-aware hyperbolic embeddings for recognizing unknown objects.
Study on LLM-based network AI agents executing network procedures via tool-calling sequences; evaluates four approaches.
Framework combining LLMs and vision-language models with adaptive control for contact-rich robotic manipulation tasks.
arXiv research on cross-view multi-object tracking using natural language with weak supervision and foundation models.
Research on emotion recognition in conversations using pre-trained language models with focus on interpretability and handling imbalanced datasets.