Same Voice, Different Lab: On the Homogenization of Frontier LLM Personalities
Large-scale analysis showing frontier LLM personalities converge toward similar trait expressions despite different training.
Large-scale analysis showing frontier LLM personalities converge toward similar trait expressions despite different training.
Survey of safety risks, attacks, and defenses in embodied AI systems operating in autonomous domains.
AI interface for content exploration when users have vague intent, between passive feeds and structured search.
User-centric evaluation of explainability methods for AI-based medical image diagnosis systems.
Investigation of CLIP embedding contributions to memorization in Stable Diffusion text-to-image models.
Analysis of how verification errors impact reinforcement learning with verifiable rewards for LLM reasoning tasks.
Vision-language model approach for video anomaly detection with interpretable reasoning and spatial localization.
Research showing guard models lose safety alignment through standard fine-tuning on benign data, affecting agentic AI pipelines.
Foundation model for cardiotocography analysis using self-supervised learning on clinical fetal monitoring data.
EvoJail uses evolutionary algorithms to generate diverse jailbreak prompts for LLMs, addressing adaptability and diversity in automated attacks.
Analyzes generalization bounds of Spiking Neural Networks through Rademacher complexity for theoretical understanding of neuromorphic computing.
DeRelayL proposes sustainable decentralized relay learning to democratize large-scale model training across resource-limited participants.
Proteo-R1 develops reasoning foundation models for de novo protein design that explicitly model functional residues and interactions.
Multi-agent framework for multimodal controversy detection in videos by modeling diverse audience perspectives without training data.
PrismAgent is a zero-shot multi-agent framework for detecting harmful content in memes through interpretable case-analysis reasoning.
Healthcare AI GYM provides comprehensive training environment for medical AI agents through multi-turn RL with clinical reasoning tasks.
Studies pass-rate rewards in critic-free RL for code generation with LLMs, addressing sparse reward signals on challenging problems.
RouteHijack demonstrates routing-aware attacks on Mixture-of-Experts LLMs by exploiting expert selection mechanisms to bypass safety alignment.
Proposes Kernel Affine Hull Machines to replace neural inference with lightweight analytical estimators for efficient query-side semantic encoding in retrieval.
Introduces techniques to detect LLM jailbreaks by tracing refusal dynamics through latent trajectories, enabling robust adversarial detection.
Reward Hacking Benchmark evaluates safety vulnerabilities in RL-trained LLM agents with tool access through multi-step tasks with shortcut opportunities.
AutoRAGTuner automates optimization of Retrieval-Augmented Generation pipelines through declarative configuration, eliminating manual tuning of architecture and hyperparameters.
Framework analyzing gradient transport in large language model pretraining across scales using five observables separating cascade dynamics and efficiency metrics.
Cross-lingual safety alignment framework for LLMs via self-distillation, addressing vulnerability in low-resource language jailbreak attacks.
ARIS: open-source autonomous research harness for multi-agent LLM collaboration with assurance mechanisms for long-horizon research workflows.
MechaRule: extract symbolic rules from LLM internals via contrastive hierarchical ablation, grounding decision logic in neuron circuits.
MedStruct-S benchmark for semi-structured information extraction from OCR clinical reports: key discovery, conditioned QA, and key-value pair extraction.
Transformer inference acceleration exploiting low-rank token activation manifold with gated subspace decomposition for language models.
ARISE: repository-level graph representation system enabling AI agents for fault localization and automated program repair with semantic precision across code dependencies.
PIIGuard: webpage-level defense mechanism against PII harvesting by browsing-enabled LLM assistants using adversarial sanitization techniques.
Choreographic programming language for designing protocols in multi-agent agentic systems handling self-interest, private information, and untrusted participants.
Analysis of community-developed LLM applications from 2025 hackathon for materials science and chemistry, categorizing emerging usage patterns across research workflows.
Economic analysis of how AI systems affect labor markets and human-provenance verification as infrastructure. Societal implications rather than technical.
Survey of confidential computing security for LLM-driven agents handling secrets, credentials, and sensitive context across tool-calling and multi-agent protocols.
Self-mined hardness approach for safety fine-tuning of LLMs, scoring prompts by model rollout harm rates to improve robustness.
MAGE framework protecting LLM agents from long-horizon attacks using shadow memory to detect multi-turn exploitation patterns.
Ortho-Hydra method using orthogonalized mixture-of-experts LoRA to prevent style bleed in multi-style diffusion transformer fine-tuning.
RLDX-1 technical report on vision-language-action models for robotic control with improved memory, motion awareness, and physical sensing.
SHIELD dataset and distilled language models for clinical text de-identification, balancing performance with enterprise deployment constraints.
Cryptographic provenance system for AI package registries to prevent dependency confusion attacks in software distribution.
DGPO algorithm for fine-grained credit assignment in RL-based LLM alignment, improving reasoning step isolation in chain-of-thought reasoning.
LLM-based agent framework for detecting anomalies in additive manufacturing 3D printing processes.
RAG applied to reasoning tasks by retrieving intermediate thinking traces instead of documents, improving math and code generation performance.
S3 framework for multimodal learning using mixture-of-experts to specialize, select, and route semantic experts for task-specific needs.
Training-free optimization for video vision-language models that reuses stable visual state to reduce redundant computation.
Pilot study evaluating multimodal LLMs for zero-shot recognition of pathological seizure movements in neurological disorder videos.
SkCC: portable skill compiler for LLM agents addressing 40% performance variation across frameworks by converting SKILL.md specs to platform-specific formats.
Neural ODE method for large-scale traffic forecasting handling continuous dynamics and discrete anomalies with truncation error guidance.
LLMs synthesize complete RL task interfaces (observations and reward functions) automating environment interface design for new tasks.
Theory-building world models inspired by developmental cognition: learning internal representations of world dynamics from observation.