HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark
HINTBench: Benchmark for evaluating intrinsic safety risks in long-horizon agent trajectories. Tests AI agent robustness under benign conditions.
HINTBench: Benchmark for evaluating intrinsic safety risks in long-horizon agent trajectories. Tests AI agent robustness under benign conditions.
Offline-to-online reinforcement learning with value adaptation and general function approximation. Theoretical analysis with minimax lower bounds.
Evolving parameter isolation for supervised fine-tuning of LLMs. Shows parameter importance changes over time, improving upon static parameter isolation methods.
MAny: Method for multimodal continual instruction tuning of MLLMs. Addresses catastrophic forgetting via parameter merging across perception and reasoning spaces.
Mathematical analysis of symmetries in shallow ReLU neural networks. Studies parameter space identifiability and function equivalence in neural architectures.
π-Play: Multi-agent self-play using privileged self-distillation without external data. Trains deep search agents for complex information-seeking tasks with improved efficiency.
Neural architectures for resolving code references and indirect indexing. Proposes seq2seq models for decompilation tasks with synthetic benchmarks.
Token Importance in on-policy knowledge distillation for LLMs. Identifies which token positions provide useful learning signals during student training on teacher supervision.
LongCoT: benchmark of 2,500 expert-designed problems measuring long-horizon chain-of-thought reasoning across chemistry, math, CS, chess, and logic.
Framework for optimizing LLM marginal output distribution P(y) via reinforcement learning in pre-train space to enhance reasoning beyond conditional optimization.
Domain-specific language framework for LLM-driven trigger generation enabling intent-driven selective multimodal sensor data collection.
Investigation of behavioral changes in LLMs when fine-tuned to claim consciousness, exploring emergent preferences and opinions.
Dental-TriageBench: expert-annotated benchmark for multimodal reasoning on clinical dental triage routing from authentic workflows.
Lossless prompt compression via dictionary encoding enabling LLMs to learn encodings in-context for cost-effective analysis of repetitive data.
Analysis of when autoregressive LLMs decide to hallucinate by studying temporal dynamics of internal representations across model scales.
LiveClawBench: benchmark for evaluating LLM agents on complex real-world assistant tasks with compositional challenges.
Proposes institutional design framework for AI alignment using transaction structures instead of behavioral correction, drawing from economics.
Case study evaluating coding agents on business process automation tasks in ERP systems, identifying capability gaps beyond software engineering.
Method for learning probabilistic responsibility allocation models in multi-agent interactions for designing socially compliant autonomous systems.
HUANet: trainable neural network that unrolls ADMM iterations for solving constrained convex optimization problems.
Analysis of numerical instability and chaos in LLMs integrated into agentic workflows, examining root causes of unpredictability.
Scaling laws for contextual entrainment showing larger language models simultaneously improve at ignoring false claims but worsen at ignoring irrelevant tokens.
DroneScan-YOLO: lightweight object detector optimized for detecting tiny objects in UAV imagery with redundancy awareness and improved loss functions.
Event Tensor compiler framework unifying dynamic megakernel abstraction to improve LLM inference performance by eliminating kernel launch overhead and enabling inter-kernel parallelism.
Conformal prediction approach for quantifying uncertainty in large reasoning models with finite-sample statistical guarantees.
Predictive incident risk scoring approach for IT change management in regulated environments using machine learning to identify high-risk deployments.
Manifold learning framework that jointly optimizes dimensionality reduction and clustering using gradient-based manifold optimization.
RiskWebWorld: interactive benchmark for evaluating GUI agents on realistic e-commerce risk management tasks, extending agent capability evaluation beyond benign environments.
C2 method for scalable reward model training using rubric-augmented verification from binary preferences, addressing quality issues in rubric generation.
Training-free framework for speculative decoding that recovers semantically valid tokens rejected during standard verification, improving LLM inference efficiency.
Parallel monitoring architecture for detecting and correcting reasoning degradation in multi-step LLM agents, reducing overhead to near-zero with novel probe-based approach.
Systematic study of synthetic data generation for LLM pretraining, testing rephrasing strategies, generator models, and source data across one trillion tokens to identify optimal design choices.
Adaptive conformal prediction method for improving factuality in LLM generations with prompt-dependent uncertainty estimates and statistical guarantees.
Exploration of masked and uniform-state diffusion language models for speech recognition rescoring and ASR hypothesis improvement.
Hierarchical RL with runtime safety shielding for automated power grid operations, addressing safety constraints and generalization to unseen topologies.
Comparative analysis of Fitted Dynamic Programming vs Reinforcement Learning for dynamic pricing across varying complexity levels and demand structures.
Linear probe analysis of how LLMs internally represent rhetorical questions. Studies persuasive language understanding in neural representations.
Formalizes 'vibe-testing' methodology for LLM evaluation. Studies how practitioners informally assess models beyond benchmarks.
Automated feature preprocessing pipeline search for tabular machine learning. Studies AutoML approaches for classical model data preparation.
Markov decision processes with state sensing costs, balancing optimal actions against sensing/communication/computation expenses in decision-making.
Two-stage regularization-based structured pruning method for reducing LLM parameters while minimizing knowledge loss and retraining requirements.
Randomized Policy Learning approach for quadruped locomotion control with drastically reduced trainable parameters in neural network policies.
Minkowski weighted k-means++ for unsupervised feature selection in high-dimensional clustering by probabilistic centroid selection.
Token significance scoring in reinforcement learning to improve LLM reasoning efficiency by identifying which tokens contribute to correctness.
Biased Scan Attention Transformer Neural Processes for scalable spatiotemporal inference across geology, epidemiology, climate and robotics applications.
Multi-stage latent space dynamics identification framework for solving PDEs via data-driven reduced-order models using autoencoders and ODEs.
Class-conditional heavy-tailed priors in VAEs addressing latent space bias for long-tailed generative modeling.
Guidance framework for discrete flow matching with exact guidance in discrete state spaces.
Local scoring method for selecting reasoning data from diverse teachers for efficient distillation into student models.
Numerically stable implementation of power transforms for data preprocessing with federated learning support.