Generative Diffusion Models of Stochastic Graph Signals
Generative diffusion models for sampling stochastic signals on graphs, applied to recommender systems and financial forecasting.
Generative diffusion models for sampling stochastic signals on graphs, applied to recommender systems and financial forecasting.
LEMUR 2: Large-scale extensible NAS benchmark with 14,000+ distinct architectures and 750,000+ training records for cross-domain neural architecture evaluation.
Empirical study using rule-based expert as benchmark to evaluate lightweight RL agents for imperfect-information card games across 100+ experiments.
On-policy self-distillation method for LLMs using geometric approaches where teacher model has access to solution hints to improve student model reasoning.
Best-arm identification algorithm that pairs costly reward observations with cheap proxy scores from LLMs to improve data-driven decision-making efficiency.
Self-supervised image clustering framework using evolutionary algorithms instead of gradient descent, eliminating need for predefined loss targets.
DNN architecture learning for edge devices using zeroed batch normalization to meet strict latency constraints in real-time applications.
Wearable foundation model using physical activity data for scalable broad-spectrum health prediction.
Federated learning approach for rapid model adaptation under real-world client churn in recommendation systems.
Unbounded Positive optimization for RL addressing exploration-stability dilemma in LLM reasoning via importance sampling.
Mechanistic analysis dissecting sycophancy in LLMs into factual and opinion subtypes using internal representations.
Study showing online data selection during fine-tuning acts as implicit alignment mechanism for LLM behavioral preferences.
Constrained decoding for diffusion language models via finite automata enabling structured outputs like JSON schemas.
Gimitest: Open-source comprehensive testing framework for single and multi-agent RL policies across varying conditions.
Interpretable ML framework for tabular data addressing feature interactions through sparse rules and patterns.
Studies adversarial robustness of relational deep learning on heterogeneous temporal graphs with integrity constraints.
K-Risk dataset combining high-risk driving scenarios with LLM annotations and semantic labels for autonomous driving safety research.
Studies causal interventions in language model components to understand task behavior, extending beyond global activation-space steering.
Proposes lossless symbolic storage for KV-cache using contractive iterated-map codes to reduce memory cost in long-context LLM inference.
Analyzes optimizer implicit bias through information allocation dynamics, explaining how training signals distribute between weights.
Studies multi-task agentic RL for LLM-based agents, identifying exploration-exploitation dynamics across different tasks.
Method for pre-deployment safety evaluation of LLMs by simulating realistic deployments from de-identified conversations to assess failure rates.
Modular language for characterizing reachable gradient methods, auditing optimizer mechanisms and their interactions systematically.
Proposes HPG-Diff, physics-guided diffusion framework for topology optimization with differentiable connectivity constraints.
Combines reinforcement learning with model predictive control to enforce hard safety constraints during exploration in cyber-physical systems.
Presents FMMVCC, a Mamba-based clustering method for unsupervised time series analysis using fuzzy logic and contrastive learning.
Analyzes memorization-based privacy attacks (TATD) in federated learning, studying how malicious training can exfiltrate data in distributed settings.
Comprehensive overview of mechanistic interpretability for reverse-engineering neural network internals, including circuit analysis and symbolic reasoning approaches.
Proposes SDE framework for uncertainty estimation in hypergraph neural networks, addressing uncertainty from higher-order relations and complex dependencies.
Studies multi-agent AI control techniques to prevent coordinated attacks across distributed AI deployments, addressing risks like model-weight exfiltration and training poisoning.
Research on adversarial vulnerability in vision-language models through spectral analysis of intermediate linear transformations.
Research benchmark evaluating open-weight vision-language models for fast radio burst detection; zero-shot generalist approach vs specialized detectors.
Research introducing Sparse Delta Memory architecture for scaling linear RNNs; improves long-context recall while reducing FLOPs.
Theoretical analysis of sample complexity for learning autoregressive chain-of-thought traces; proves bounds governed by local next-token classification.
Research on reinforcement learning for real-world agents with irreversible interactions; proposes penalizing unsafe paths while rewarding outcomes.
Research benchmark studying interaction between differential privacy and fairness-aware learning on synthetic tabular data in high-stakes ML.
Research: FFT-based spectral preprocessing of query-key projections improves transformer attention; 79% validation loss reduction on TinyShakespeare.
RAID framework uses reward-adaptive RL for automated game testing, reducing retesting effort after behavior modifications in NHL26 development via iterative goalie AI exploit discovery.
TimEE applies in-context learning to end-to-end time series classification, replacing two-stage train-then-classify pipeline with unified learning from label demonstrations.
Asynchronous single-rollout RL system for LLM post-training on long-horizon agentic tasks, improving efficiency and training stability over synchronous batch-interleaved approaches.
Theoretical explanation for self-supervised learning efficiency: data augmentation induces graph structure enabling fast transductive rates O(1/n_L) with graph-Laplacian regularization.
Collaborative synthetic data generation for one-shot federated learning, enabling knowledge transfer across divergent client distributions in single communication round.
Compares multi-class vs multi-label BERT formulations for CVE-to-CWE vulnerability mapping, analyzing how taxonomy structure shapes classification errors.
Formulates neural network depth adaptation as optimal control problem with posteriori error estimation, enabling principled layer insertion based on approximation error distribution.
Analyzes classifier-free guidance breakdown in diffusion models through numerical analysis, proposing terminal-fitted repair to stabilize sampling at high guidance scales.
Analyzes how RoPE frequency usage in transformers matches training data's relative-distance structure, with implications for length generalization and position embeddings.
NOTES integrates neural operators with evolutionary strategies for PDE-constrained inverse design optimization, combining robustness with high-dimensional search capability.
Novel method combining selective timestep weighting and advantage-based replay to improve sample efficiency of RLHF applied to diffusion models, reducing feedback requirements.
Shows trusted monitoring models overfit to specific untrusted policy families and don't generalize across model lineages.
HiFuzz uses hierarchical reinforcement learning with structured generation agents for CPU fuzzing and processor verification.