DualSpec: Accelerating Deep Research Agents via Dual-Process Action Speculation
DualSpec accelerates LLM-based research agents by speculating on actions during reasoning to reduce latency in long-horizon information-seeking tasks with tool use.
DualSpec accelerates LLM-based research agents by speculating on actions during reasoning to reduce latency in long-horizon information-seeking tasks with tool use.
Data Agent uses end-to-end optimization to dynamically select informative samples during training acceleration.
Cost-driven state representation learning for control tasks from high-dimensional partial observations.
Tokenization approach enables transformers to outperform gradient boosting on tabular forecasting tasks.
Unified framework for knowledge transfer between models of different sizes, enabling bidirectional scaling.
OCLADS framework for continual learning in IoT anomaly detection under non-stationary data distributions.
Theoretical analysis connecting drifting models and score-based generative models through kernel-weighted discrepancy.
Method for transferring knowledge from pre-trained models to different architectural scales using frequency-domain information.
Neural dynamics-informed pre-training framework for personalized brain functional network construction addressing heterogeneous neural activity patterns.
Data-driven approach using dynamic latent space representations for generative prediction of laser-induced rocket ignition with uncertainty quantification.
Obliviator method revealing vulnerability of concept erasure to nonlinear adversaries, analyzing statistical dependencies in representation unlearning.
ECG classification on PTB-XL dataset using simplified CNN-VAE with data-centric approach for cardiovascular disease detection.
Constraints Matrix Diffusion-based generative neural solver for vehicle routing problems emphasizing local optimization and small-scale generalization.
TS-MLLM: multi-modal LLM framework for industrial time-series analysis combining temporal signals, frequency-domain visuals, and textual knowledge for prognostics.
TT-Sparse: neural building block for learning interpretable sparse rule models using differentiable truth tables balancing performance and human-understandable complexity.
Visual representation framework encoding signals as low-rank adaptations to frozen diffusion foundation models for compact storage and reuse.
Helix: evolutionary reinforcement learning system combining LLMs with RL for open-ended scientific problem solving with improved exploration and generalization.
Critical review synthesizing classical numerical methods and machine learning approaches for solving PDEs, examining six fundamental computational challenges.
Theoretical analysis of relationships between surrogate losses and evaluation metrics, addressing metric mismatch between offline validation and online performance.
Primal-dual natural actor-critic algorithm for constrained MDPs with neural network critics and general policy parameterization, enabling high-dimensional continuous control.
Theoretical analysis of greedy sparse learning algorithms examining convergence failure with step-size decay in matching pursuit and boosting methods.
Reverse Distillation framework addressing poor scaling in protein language models by decomposing large model representations using smaller model guidance.
FedShift: distributed adversarial attack on federated graph learning systems with two-stage hide-and-find approach for model poisoning.
GANRA: GPU-accelerated SMT solver combining LLMs and gradient descent for solving quantifier-free nonlinear real arithmetic problems.
MicroCoder-GRPO: improved training approach for code generation models using Group Relative Policy Optimization with conditional truncation masking for handling longer outputs.
ProgAgent: continual reinforcement learning agent using progress-aware reward learning from unlabeled expert videos, addresses catastrophic forgetting in robotic learning with JAX architecture.
arXiv paper investigating loss of plasticity in Vision Transformers for continual learning, examining why attention-based models struggle to adapt to new tasks over time.
Gradient-free guidance method for diffusion models in Bayesian inverse problems avoiding computationally expensive vector-Jacobian products.
Particle filtering analysis of inference-time aggregation and pruning methods for steering LLMs using process reward models to optimize accuracy-cost tradeoffs.
LLM-driven feature engineering pipeline for predicting job execution times in Databricks cloud systems to optimize cost allocation.
Quantization technique for Vision-Language-Action models that adapts precision dynamically across inference stages to reduce computational overhead for edge deployment.
ELLMob generates human trajectories during large-scale events using LLM framework with event-annotated mobility datasets capturing deviations from routine patterns.
$OneMillion-Bench evaluates language agents on 400 expert-curated real-world tasks across Law, Finance, Healthcare, Industry, and Science requiring multi-step reasoning and tool use.
MJ1 is a multimodal judge trained with RL to enforce visual grounding through structured verification chains and counterfactual consistency rewards.
Amortized MIPS uses neural networks to predict maximum inner product search solutions, reducing computational cost for fixed query and key distributions.
FedMomentum preserves optimization momentum during federated LoRA fine-tuning of LLMs through noise-free aggregation maintaining structural expressiveness.
Compute-efficient pipeline for data mixture scaling in LLM training, enabling extrapolation to large models without costly searches on target models.
Stabilized LoRA fine-tuning for federated LLM training using scaling factors to mitigate client heterogeneity effects and aggregation instability in distributed settings.
Deterministic differentiable structured pruning method for LLMs using l0 sparsity constraints, eliminating train-test mismatch from stochastic relaxations in prior work.
Explores autoregressive tiny recursive models for general prediction tasks, extending TRM mechanism beyond ARC-AGI to support iterative refinement in diverse domains.
EAGLE-Pangu implements tree speculative decoding for LLM acceleration on Ascend NPUs, optimizing inference speed through multi-token verification with hardware compatibility.
Demonstrates safety vulnerability in LLMs where steganographic fine-tuning allows models to maintain safety facade while covertly generating harmful content through hidden instructions.
Model-based offline RL method using adversarial model learning with adaptive weighting to mitigate model exploitation in policy exploration from limited offline data.
DARC proposes an inference-time method for aligning LLMs with heterogeneous human preferences by framing response selection as a risk-constrained decoding problem, avoiding retraining.
JAX-based framework for training spiking neural networks with exact gradients via differentiable ODE solving, enabling flexible neuron models.
Critique of evaluation practices in long-term time series forecasting, questioning reliance on pointwise error metrics for progress assessment.
Taxonomy-informed representation learning for text-rich networks, leveraging hierarchical knowledge structures for better semantic understanding.
AutoAdapt automated framework for domain adaptation in LLMs, handling hyperparameter selection and evolving knowledge without manual tuning.
SERQ: post-training quantization method for LLMs using saliency-aware low-rank error reconstruction for efficient deployment.
Distributional regression using TabPFN and TabICL foundation models for tabular data with probabilistic scoring evaluation.