Leveraging Human Feedback for Semantically-Relevant Skill Discovery
Skill discovery in RL using human preference feedback to ensure safe and aligned agent behaviors.
Skill discovery in RL using human preference feedback to ensure safe and aligned agent behaviors.
Meta-Aligner optimizes multi-objective LLM alignment by dynamically adjusting preference weights during training.
BitRL enables RL agents on edge devices by quantizing LLMs to 1-bit, reducing memory and computational requirements.
Self-Abstraction Learning framework addressing gradient vanishing, overfitting, and training stability in deep neural networks.
Methods to mitigate catastrophic overfitting and error amplification in Fast Adversarial Training.
Research on catastrophic overfitting in Fast Adversarial Training, analyzing backdoor mechanisms in adversarial robustness.
Unified plugin framework for composable controllable diffusion methods across different backbones and tasks with compatible training pipelines.
Pilot-Activated Recovery System using soft-actor critic reinforcement learning with hyperparameter optimization for aircraft upset recovery.
DPRM: Token-ordering module for diffusion language models using Doob h-transform to improve generation efficiency and exploration.
Analysis of linear region complexity in self-supervised ReLU networks during training, extending prior supervised learning research.
AI-based Automatic Ground Collision Avoidance System using reinforcement learning for advanced jet trainers.
Pretrained molecular embedding distance for ligand-based virtual screening and goal-directed molecular generation in drug discovery.
SceneSelect: Selective learning framework with expert scheduling for trajectory prediction across heterogeneous scene types.
Multi-objective reinforcement learning using reward-free RL perspective to train policies adapting to different user preference weightings.
Global optimization algorithm for noisy function evaluation requiring only weak local smoothness assumptions without prior knowledge.
GradMAP: Decentralized multi-agent proximal learning for coordinating grid-edge devices while respecting AC distribution network constraints.
Online learning algorithm for bandit problems with partial observability, where learners observe losses of other actions beyond their own.
Spatio-temporal graph neural networks for detecting market manipulation and fraud in cryptocurrency transactions.
Unsupervised machine learning clustering analysis of social media usage patterns and mental health correlations.
Functional Task Networks (FTN) for continual learning using parameter isolation inspired by mammalian neocortex structure, avoiding catastrophic forgetting.
Proposes agent-native research artifacts to address limitations of traditional scientific papers, enabling AI agents to access experimental branching and engineering details.
Physics-informed framework for feature selection in high-dimensional data using diffusion and spectral embedding without greedy search.
Method for training large neural networks by using GPU replicas to explore multiple learning rates simultaneously with minimal communication overhead.
Benchmark for evaluating generalization of specification-guided reinforcement learning agents across diverse environments and unseen formal specifications.
Research on learning from Chain-of-Thought supervision across multiple different solution approaches to improve model generalization on tasks like math problems.
arXiv report reviewing AGI forecasting methodologies, their limitations, and gaps in current prediction approaches.
Systematic study measuring intrinsic non-randomness in LLM token distributions via Entropic Deviation metric across models and prompts.
arXiv paper on evaluating VLM performance on multi-line handwritten math OCR with semantic reasoning metrics.
arXiv paper on intelligent fault diagnosis for aircraft using multi-fidelity digital twin and LLM-enhanced interpretable reporting.
arXiv paper on generative self-supervised learning framework for physiological parameter estimation from photoplethysmography data.
arXiv paper on accelerating RL training for wind farm control using domain knowledge from steady-state wake models.
arXiv paper on multi-agent reinforcement learning for wind farm flow control with structural load constraints using I-SAC architecture.
arXiv paper combining reinforcement learning with model predictive control for wind farm wake steering optimization.
arXiv paper on deep learning for RF interference rejection in signal detection and demodulation across varying SINR levels.
Speech-aware LLM extended with word-level timestamp prediction for ASR applications like captioning and media synchronization.
Structural model analyzing how AI trading agents with similar representations create systemic instability in financial markets.
Audio2Tool dataset for evaluating tool-calling capabilities of speech language models across smart home, finance, and navigation domains.
Browser-based tool for training CNN models on microcontrollers with image collection, training, and deployment pipeline for edge devices.
LLM-based agent system for fine-grained information retrieval grounded in scientific literature for research queries.
Framework addressing context-fragmented violations in multi-agent systems where distributed agent actions collectively violate policies.
Framework for learning compact executable verifiers for LLM outputs combining interpretability with expressive capability.
LLM-powered pipeline for automated extraction and structuring of materials science data from scientific literature.
Variable-step diffusion model framework for efficient medical image translation with reduced computational cost.
Improved trust region Bayesian optimization strategy addressing lengthscale design issues in high-dimensional settings.
3D asset dataset of 10,000+ spatially and semantically aligned objects for embodied AI and robotics simulation.
Self-supervised learning framework for Android malware detection addressing temporal bias in training data.
Framework teaching LLMs biomedical reasoning through counterfactual imagining for clinical trial outcome prediction.
Transformer-based framework for causal inference from observational data with improved handling of complex treatment mechanisms.
Mixture-of-Experts architecture with heterogeneous expert sizes for improved LLM scaling and efficiency.
Formal language learning perspective comparing fine-tuning vs in-context learning in LLMs with controlled experimental setup.