Research on distributed vs on-device inference tradeoffs for DNNs in cyber-physical systems, examining network latency, energy consumption, and real-time control constraints.
Federated learning framework enabling concurrent multi-task training across heterogeneous decentralized devices in privacy-preserving manner.
Framework combining human expertise with Bayesian optimization for accelerated discovery in data-scarce scientific domains like fusion energy.
Kisan AI: profit-aware crop advisory system combining ML for yield prediction with economic analysis to optimize farmer financial outcomes.
ARHQ: post-training quantization method for low-bit LLM weight quantization using residual Hessian analysis to mitigate error propagation.
Wasserstein distributionally robust approach for RLHF addressing reward misspecification in LLM alignment and policy optimization.
Consistent diffusion language models: method to accelerate discrete diffusion models for language generation via consistency training adaptation.
Study analyzing how supervised fine-tuning affects generative diversity in large language models through formal empirical testing.
State Stream Transformer V2: transformer variant with FFN-driven nonlinear recurrence for parameter-efficient reasoning in continuous latent space.
Study showing advanced jailbreaks of frontier LLMs impose negligible performance degradation, with complexity inversely scaling with model capability across Claude variants.
Caracal proposes O(L log L) Multi-Head Fourier module to replace attention in LLMs, addressing quadratic scaling and positional encoding limitations via FFT-based sequence mixing.
ArXiv: Caracal LLM architecture replaces quadratic attention with O(L log L) Multi-Head Fourier module for efficient long sequence processing.
ArXiv: Data Deletion in Adaptive RL studies context learning in time-varying environments via contextual Markov Decision Processes.
ArXiv: Federated learning framework for weather modeling using distributed sensor data without sharing raw data.
ArXiv: Conformalized Quantum DeepONet Ensembles reduce operator learning complexity and improve uncertainty quantification for high-dimensional systems.
ArXiv: Odysseus scales vision-language models to 100+ turn decision-making in games using reinforcement learning, extending beyond short-horizon settings.
ArXiv: HyperODE RCA combines hypergraph learning and latent ODEs for root cause analysis in microservice systems.
ArXiv: Binomial flows framework for denoising and flow matching on discrete ordinal data using Tweedie's formula.
ArXiv: Uniform-Correct Policy Optimization addresses diversity collapse in reinforcement learning with verifiable rewards on reasoning tasks.
ArXiv: AlphaInventory uses LLMs for evolutionary search to optimize inventory policies in dynamic online environments with deployment guarantees.
Reinforcement learning method improving LLM reasoning via negative sample projection, boosting diversity in reasoning tasks.
Reframes LLM ensembling as mixture model optimization problem to improve performance while reducing computational cost.
Post-training quantization method for binarizing both weights and activations in LLMs, enabling end-to-end acceleration.
Demonstrates LLM vulnerability to row/column permutations in tabular data, revealing robustness issues in Table Question Answering tasks.
Federated sketch contextual linear bandits reducing computation and communication costs via SVD sketching for high-dimensional data.
Scale-aware adversarial analysis diagnostic for evaluating whether generative AI internalizes physical laws in multiscale complex systems.
Hierarchical abstract tree indexing for cross-document multi-hop retrieval-augmented generation with improved distribution adaptability.
Stable-GFlowNet for diverse and robust LLM red-teaming using contrastive trajectory balance and generative flow networks.
Possibilistic framework for epistemic uncertainty modeling in deep neural networks balancing principled theory with computational efficiency.
Improving sparse MoE routing at domain transitions with lightweight gate modifications using temporal memory and entropy regularization.
Task vector approach for integrating SFT and RLVR LLM post-training paradigms to combine knowledge breadth and reasoning depth.
First study of machine unlearning for offline stochastic multi-armed bandits with privacy implications for sequential decision systems.
Research on budget-constrained optimization for ML problems like quantization and pruning using manifold geometry and evolutionary search.
Adam-style zeroth-order optimizer for memory-efficient LLM fine-tuning using only forward passes.
Risk-averse Q-learning method using Markov coherent risk measures and multipattern approximation with regret bounds.
Augmented Lagrangian multiplier network for enforcing state-wise safety constraints in reinforcement learning.
Evaluation framework testing LLM reasoning capabilities via formal proof synthesis in alien mathematical domains.
Compositional graph embedding method using Aitchison geometry for interpretable node representation learning.
Theoretical extension of Weisfeiler-Lehman test to combinatorial complexes for topological neural network expressiveness.
Multi-agent Monte Carlo tree search optimization using interaction-guided exploration for cooperative game playing.
Multi-agent platform for executing natural-language plans with constraint-guided LLM execution and control constructs.
LLM-based workflow with validation for generating executable statistical chart code from tabular data.
Reinforcement finetuning approach for time series foundation models to improve generalization under distribution shifts.
AI agents for autoformalizing memory specifications in chip design verification, translating specifications to formal representations.
Being-H0.7: Latent world-action model from egocentric videos for robot control using visual-language-action representations.
DeGenTWeb: Analysis of LLM-generated content prevalence on web, finding LLM detectors perform worse than advertised.
Quantum Gaussian processes: Bayesian framework for learning from quantum systems with interpretability and scalability.
Mutation testing methodology for quantum machine learning models to verify implementation correctness and detect faults.
ViLegalNLI: First Vietnamese legal domain NLI dataset with 42,012 premise-hypothesis pairs for legal reasoning.
Adaptive norm-based regularization strategies extending ridge and lasso penalties to neural networks using input covariance structure.