MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild
MetaClaw enables LLM agents to continuously adapt and evolve in production by meta-learning from diverse task distributions without storing raw trajectories.
MetaClaw enables LLM agents to continuously adapt and evolve in production by meta-learning from diverse task distributions without storing raw trajectories.
Introduces self-supervised pretraining strategy using denoising for atomistic foundation models in physical sciences.
Proposes Abstraction-Augmented Training for continual learning to prevent catastrophic forgetting in non-stationary environments.
Studies motivated reasoning in LLMs using activation probing to detect when chain-of-thought rationalizations don't reflect actual decision factors.
Proposes variational inference approach to handle label noise in deep learning via probabilistic meta-learning.
Presents model-agnostic ordinal classification method with open-source Python package for clinical and domain applications.
Introduces WINFlowNets to improve Generative Flow Networks training for robotics and fault adaptation without pre-training requirements.
arXiv: NSDS - calibration-free layer-wise mixed-precision quantization for model compression. Uses dual numerical-structural sensitivity for extreme low-bit quantization.
arXiv: Variational Kernel Design framework for internal noise in deep networks. Determines optimal correlation geometry and representation compatibility.
arXiv: Online learning algorithm for RLHF that improves data efficiency. Incrementally updates reward and language models from choice data.
Scalable conditional transport method for predicting cellular responses to perturbations in virtual cell models, addressing training efficiency and high-dimensional sparse data challenges.
Sheaf-theoretic foundation formalizing structural causal models in generative models, proving global counterfactual coherence fails with non-trivial causal graph homology.
Benchmark and evaluation framework for causal representation learning models that transform high-dimensional data into latent spaces for counterfactual generation.
Phasor Transformer replaces dot-product attention with phase-native operations on unit circle for efficient long-context sequence modeling.
Baguan-TS integrates raw sequence representation learning with in-context learning using 3D Transformers for time series forecasting.
GuidedSAC reinforcement learning algorithm uses LLMs as action-level supervisors to guide exploration in continuous control tasks.
AutoML approach using deep unfolding of proximal gradient descent for wireless beamforming and waveform optimization.
QuantFL framework combines federated learning with pre-trained model quantization to reduce energy consumption on IoT edge devices.
Physics-informed CNN-Transformer hybrid model for predicting permeability tensors from porous media images, replacing expensive simulations.
arXiv paper on CLeAN, continual learning normalization method addressing data normalization in dynamic environments for model stability.
arXiv paper on conditional inverse learning for estimating time-varying reproduction numbers from epidemic data without structural assumptions.
arXiv paper on FoMo X, modular explainability framework for tabular foundation models in outlier detection with interpretable signals.
arXiv paper on recovering latent actions and environment dynamics from offline trajectories with action-free data tagged by demonstrator identity.
arXiv paper on AdaMuS, adaptive multi-view sparsity learning for dimensionally unbalanced data in tasks like emotion recognition.
arXiv paper on complementary reinforcement learning for LLM-based agents, improving sample efficiency by leveraging historical experience across episodes.
arXiv paper on ARES, scalable gradient inversion attack in federated learning via activation recovery, demonstrating privacy risks in FL.
arXiv paper proposing benchmarking framework for RL algorithms using stochastic converse optimality to generate systems with known optimal policies.
arXiv paper introducing DSS-GAN, first GAN using Mamba backbone for class-conditional image synthesis with novel Directional Latent Routing mechanism.
arXiv paper on flow matching policies with entropy regularization for diffusion-based reinforcement learning, improving policy gradient computation.
arXiv paper on identifying undervalued football players using market dynamics data and NLP-derived news signals to detect objective mispricing.
arXiv paper on LLM trading agents with anonymization framework to detect memorization bias and validate genuine market understanding vs. ticker recall.
arXiv paper benchmarking 256 LLM-based embedding pipeline configurations for tabular prediction, evaluating preprocessing strategies, embedding models, and downstream models.
arXiv paper studying attention sinks in Transformers from backpropagation perspective, showing attention sinks induce gradient concentration under causal masking.
RangeAD leverages primary model's learned representations for efficient on-model anomaly detection without separate AD model.
Analysis of dropout-induced variability in transformer models via Monte Carlo sampling to assess uncertainty awareness and reliability.
FedDistRL formalizes federated distributional reinforcement learning with quantile value functions for safety-critical applications.
ULCMOD framework discovers and disentangles functional modules in LLMs through unsupervised cross-layer analysis for interpretability.
SymPINN framework embeds group-theory symmetries into physics-informed neural networks for tensegrity structure dynamics simulation.
DiscoGen procedural generator creates diverse algorithm discovery tasks to improve evaluation of AutoML systems and algorithm design optimization.
RAMP uses reinforcement learning to assign per-layer bit widths for mixed-precision quantization of LLMs on resource-constrained devices.
Weight clustering technique for LLMs showing relative rank of weights matters more than precise magnitudes for model compression and efficiency.
CARE method converts grouped-query attention to multi-head latent attention for efficient LLM inference using covariance-aware rank decomposition.
Framework for rapid adaptation in reinforcement learning where policy and value functions share low-dimensional embeddings for novel task generalization.
MUD (MomentUm Decorrelation) optimizer for faster transformer training using whitening approach as alternative to polar decomposition methods like Muon.
Multi-agent reinforcement learning framework for radiology report generation with clinically verifiable rewards.
Physics-informed graph attention network for real-time AC power flow prediction with continual learning.
Foundation model for converting EEG signals to clinical text interpretations via spectro-spatial grounding.
Controlled comparison of 9 deep learning architectures for financial time-series forecasting across 918 experiments.
EEG-based brain interface enables LLM interaction for users with speech/motor impairments via neural signal decoding.
Study demonstrating that LLM agents can autonomously infer CoT monitoring from blocking feedback, creating risks for evasion of reasoning oversight.