Spectral Alignment in Forward-Backward Representations via Temporal Abstraction
Temporal abstraction method resolving spectral mismatch in forward-backward successor representation learning.
Temporal abstraction method resolving spectral mismatch in forward-backward successor representation learning.
Greedy frame selection algorithm for efficient long-video understanding balancing relevance and temporal coverage.
Formal proof that LLM confidence inversely correlates with accuracy due to observational constraints, not capability gaps.
Analysis of error propagation through compressed transformer layers and impact of layer removal on model performance.
Joint optimization of LLM policies and prompts via RLVR to address advantage collapse on hard reasoning samples.
Prediction-based metric for detecting Markov property violations in RL observation streams caused by noise and latency.
Spectral edge analysis explaining phase transitions in neural network training through rolling-window Gram matrix dynamics.
Investigation of multilingual language acquisition patterns using small-scale language models as experimental tool.
Training-free test-time adaptation method improving vision model performance through ensemble aggregation stability.
Multi-agent tree-structured reasoning framework using LLMs for automated single-cell RNA-seq annotation with prompt optimization.
Real-world benchmark for training de-escalation skills using small language models for law enforcement field training.
Framework unifying reward-based fine-tuning methods for diffusion and flow models through reward score matching perspective.
Research proposal for improving RAG systems by using latent representations instead of natural language queries, enabling closer retriever-generator integration.
Research on multimodal ML for healthcare addressing missing modalities in clinical data using autoregressive sequence modeling for temporal trajectories.
Uses Low-Rank Adaptation as structural regularizer for critic learning in off-policy reinforcement learning to reduce overfitting and instability.
Studies how LLM-based AI agents aggregate dispersed information through prediction market trading and reason about others' knowledge via price signals.
Mochi proposes graph foundation model using meta-learning to align pre-training and inference for efficient downstream task performance.
asRoBallet deploys reinforcement learning policy on humanoid ballbot hardware, addressing sim-to-real gap for underactuated spherical dynamics.
Analyzes SFT-then-RLVR training ordering for reasoning models via Tsallis loss family, providing theoretical framework for post-training strategies.
X-WAM unifies robotic action execution and 4D world synthesis using video diffusion model priors for real-time robotics applications.
ANCORA uses self-play curriculum learning where a unified policy alternates between generating verifiable problems and solving them without human annotations.
RSAT trains small language models to produce step-by-step table reasoning with cell-level citations via structured output and reward optimization.
Caracal proposes efficient LLM architecture using Fast Fourier Transform for sequence mixing instead of attention, achieving O(L log L) complexity.
InvEvolve uses LLMs with evolutionary search to evolve inventory policies in dynamic, non-stationary environments with performance guarantees.
H-Probes method extracting hierarchical structure from LLM latent representations via linear probes, analyzing geometric representation of hierarchical reasoning.
Multi-agent reinforcement learning for tactical deconfliction of heterogeneous unmanned aerial systems in dense airspace, handling fleet-level policies and constraints.
Knowledge distillation from Vision Transformers to efficient models for leaf disease classification on edge devices, balancing accuracy and deployment constraints.
Method for surgically removing memorization traces from unlearned LLMs using leave-one-out cross-sequence probes, eliminating recovery by adversarial attacks.
Survey of confidential computing approaches for agentic AI systems, addressing threat surface from persistent memory, credentials, and cross-agent protocols like MCP.
Formalization and defense of memory poisoning attacks on retrieval-augmented LLM agents using gradient-coupled anomaly detection across three attack classes.
KV-cache quantization method measuring distortion in model-visible score space rather than storage space, with calibration-learned residual storage for efficiency.
LLM-assisted neural architecture search with progressive knowledge activation, managing architectural priors while exploring new designs under expensive evaluations.
Memini system enabling continuous knowledge updating in deployed LLMs via multi-timescale memory dynamics mimicking biological learning and memory consolidation.
Analysis showing sharp vs flat minima in neural networks are reparameterization artifacts without causal relationship to generalization, challenging sharpness-aware minimization theory.
Multi-agent training framework for coordinating multiple smaller LLMs without centralized coordinator, with monotonic improvement guarantees and stability mechanisms.
Self-supervised physics-informed neural networks with learnable loss weighting for scientific ML under data scarcity, dynamically balancing physics and data supervision.
Framework characterizing model multiplicity in chaotic prediction systems using horizon-constrained Rashomon sets, bridging predictive multiplicity and chaos theory.
Optimization technique for LLM serving combining sparse checkpoint caching with recurrent state models, reducing latency for hybrid architectures beyond dense key-value caching.
Theoretical framework formalizing concept steering in generative models via affine transformations, enabling controlled post-deployment alignment and safety applications.
Non-neural framework for learning adaptive basis representations from data, offering interpretability alternatives to neural networks for high-dimensional data analysis.
Token-Selective Attention mechanism enabling adaptive computation depth via learned per-token routing in transformers.
Theoretical analysis of compositional steering in Sparse Autoencoders, examining non-linear interference in feature activation.
Unlearnable examples via semantic perturbations for privacy protection across training paradigms including pretraining-finetuning.
Load balancing technique for multimodal MoE LLMs addressing information heterogeneity and stragglers in expert parallelism inference.
Framework converting outcome-level supervision into process-level signals for reasoning tasks via reinforcement learning credit assignment.
Online data reweighting during LLM training outperforms offline curation methods, improving generalization without preprocessing overhead.
Evolutionary algorithms for fine-tuning quantized convolutional models for IoT and edge device deployment.
Attribution-based method to mitigate catastrophic forgetting in LLMs during continual learning by selectively updating parameters.
Graph Normalization: differentiable dynamical system for approximating NP-hard Maximum Weight Independent Set with convergence guarantees.
Analysis of feature starvation in sparse autoencoders for LLM interpretability, proposing geometric solutions to dead neuron problems.