Evaluating Feature Dependent Noise in Preference-based Reinforcement Learning
Evaluates feature-dependent noise in preference-based reinforcement learning with realistic noise patterns correlated to observations.
Evaluates feature-dependent noise in preference-based reinforcement learning with realistic noise patterns correlated to observations.
Proposes GIFT method reconciling SFT and RL post-training for Large Reasoning Models via Gibbs initialization to prevent distributional collapse.
Solves constrained optimization problems via gradient-based methods using hierarchical score-matching spaces to overcome local optima.
Proposes neural characteristic function approach for graph domain adaptation addressing distributional shifts without manual feature design.
Studies sequential prediction with option to abstain in semi-adversarial settings mixing adversarial and stochastic instances.
Creates deep surrogate model for blast wave prediction that generalizes to out-of-distribution urban scenarios using machine learning.
Develops federated causal representation learning for decentralized counterfactual reasoning across coupled industrial systems while preserving data privacy.
Proposes Hyperparameter Trajectory Inference to adjust neural network hyperparameters post-deployment without full retraining using optimal transport.
Studies how pretrained Vision-Language-Action models resist catastrophic forgetting during continual learning in robot policy training.
Reduces transformer KV cache by using low-dimensional keys for attention selection while maintaining high-dimensional values, achieving O(log N) dimensional compression.
JAWS improves neural PDE solvers' long-term rollouts using spatially-adaptive Jacobian regularization to prevent spectral blow-up and unphysical divergence.
Adaptive channel pruning technique reduces communication overhead in split learning by selectively transmitting intermediate feature representations.
MR-Search proposes meta-reinforcement learning with self-reflection for agentic search, enabling agents to adapt strategies across episodes and improve in-context exploration.
Method for embodied agents to autonomously discover symmetry group structure for disentangled representation learning without requiring prior knowledge of group properties.
Theoretical investigation of deep residual networks' approximation capacity in continuous dynamical systems, quantifying minimal time-horizons for diffeomorphism approximation.
OMNIFLOW is a multimodal agent combining LLMs with physics-grounded reasoning for scientific tasks involving PDEs, addressing hallucinations through cross-domain generalization.
PhasorFlow: open-source Python library for computing on unit circle using complex phasors and unitary wave interference gates.
QFT: quantization-based approach for full-parameter fine-tuning of large language models with limited computational resources.
Byte-token enhanced language models for temporal point processes analysis to model event sequences with temporal dynamics and textual descriptions.
Method for improving mathematical reasoning in smaller LLMs by integrating arithmetic learning with knowledge distillation and data augmentation.
Survey of edge-cloud collaborative computing paradigms for distributed AI deployment, covering model optimization and LLM inference strategies.
Framework for constraint learning using pruned neural networks as tractable surrogates in optimization problems.
Information Imbalance metric for analyzing semantic information alignment in deep representations across text and image models.
BiomedSQL benchmark for evaluating text-to-SQL systems on biomedical knowledge bases requiring implicit domain reasoning and scientific understanding.
Study on how prompt variability affects LLM code generation quality and functionality across different user backgrounds and expertise levels.
ToolRegistry: protocol-agnostic tool management library for function-calling LLMs, addressing fragmentation in tool integration.
Analysis of GNN generalization error to explain performance variance and benchmark skew in graph neural networks.
Benchmark study showing large multimodal models fail at inductive physical reasoning beyond training distribution.
EdiVal-Agent framework for automated, fine-grained evaluation of multi-turn image editing using object-centric assessment.
Detection methods for data contamination in RL post-training phase of LLMs, addressing evaluation validity gap.
CBF-RL integrates Control Barrier Functions into RL training to enforce safety constraints during policy learning.
Neighbor GRPO extends Group Relative Policy Optimization to flow matching models with contrastive ODE-based approach for generative model alignment.
Knowledge Immunization Framework for selective knowledge erasure from LLMs via representation-aware activation signatures, addressing GDPR and safety.
Study demonstrating reward-free backdoor attacks on RL agents through compromised simulators.
Knowledge distillation framework for fine-grained visual classification using vision-language models with prompt-aware calibration.
Learnable Gaussian sampling method for inference-time scaling in latent reasoning models to improve reasoning path generation.
Method for achieving fairness in AI systems without demographic attributes for human-centered applications.
Method for fine-tuning diffusion policies with reinforcement learning for humanoid robot loco-manipulation tasks.
Pruning method for efficient large vision-language model inference by exploiting attention patterns and addressing token redundancy.
Framework for aggregating noisy heterogeneous evidence in probabilistic reasoning tasks with explicit uncertainty quantification.
Theoretical framework for aggregating multiple evidence sources in probabilistic prediction with formal guarantees for multi-evidence reasoning.
Kubernetes multi-tenancy challenges with AI agents requiring ephemeral environments. Infrastructure scaling issues.
Analysis arguing software won't become disposable despite AI coding agents, critiquing concepts like 'vibe coding' and ephemeral apps.
South Korea's SDT opens first commercial quantum-AI hybrid data center in Seoul with 20-qubit Kreo quantum computer and Nvidia DGX B200 integration.
Scheduled: Open-source AI agent integrated with Gmail that autonomously reads meeting request emails, checks calendar availability, and drafts proposed times.
ATO: GUI control panel managing multiple LLM agents (Claude Code, Codex, OpenClaw, Hermes) with workflow orchestration and MCP integration.
LA County courts pilot AI tool (Learned Hand) to summarize legal motions and draft rulings based on judge writing styles.
Conceptual framework on how AI agents transform organizational structure and decision-making beyond efficiency gains.
Case study: developer maintaining open-source Chrome extension with AI assistance for bug fixes and feature development.
OpenAI acquires Astral, integrating open source Python developer tools (uv, Ruff) into Codex ecosystem to enhance Python development tooling.