Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
Offline reinforcement learning with parametric policies under general function approximation beyond state-wise mirror descent.
Offline reinforcement learning with parametric policies under general function approximation beyond state-wise mirror descent.
Federated learning algorithm addressing statistical heterogeneity and non-IID data with proximal-balanced scaling for privacy-preserving training.
Sample-efficient hypergradient estimation for decentralized bi-level reinforcement learning in strategic decision-making and environment design.
Masked discrete diffusion model with self-aware Markov transition kernels enabling adaptive reasoning and error correction in discrete tasks.
Stable end-to-end joint embedding predictive architecture learning world models from raw pixels without representation collapse.
Multi-scale convolutional architectures for time series classification using diverse input representations and multi-representation learning.
Theoretical framework for population-based neural network training combining fast within-model optimization with slower population-level adaptation.
Multi-task supervised fine-tuning algorithm addressing heterogeneous overfitting across dataset mixtures with overfitting-aware data allocation.
Precipitation nowcasting model combining radar observations with weather foundation model priors to improve long-lead forecasting accuracy.
Analysis of systematic biases in Chinchilla scaling law fitting method applied to LLM training, showing parameter allocation errors in compute-optimal estimates.
Cloud-edge collaborative system for photovoltaic power forecasting using large models with latency constraints and robustness to weather distribution shifts.
Method for routing prompts to optimal LLMs/generative models using diversity-aware adaptive selection beyond fidelity scores.
Survey on enterprise financial risk prediction using big data and LLMs, covering AI/computer science approaches to finance and management risk analysis.
Theoretical study of feature learning in Leaky ResNets via Hamiltonian mechanics. Analyzes representation geodesics and bottleneck structures in infinite-depth limits.
Set2Seq Transformer for temporal multiple-instance learning with permutation-invariant set representations. Models internal structure and temporal relationships across timesteps.
Coded computing schemes for distributed systems with probabilistic stragglers. Extends exact computation frameworks to handle approximate recovery scenarios.
Causal framework for evaluating LLMs controlling for randomization in token generation. Proposes coupled generation model for fair model comparison and ranking.
Framework integrating ML prediction uncertainty into online algorithm design. Uses calibration to leverage prediction-level confidence in algorithms with predictions.
Gen-C: Generative framework for simulating high-level crowd behaviors in virtual environments. Captures agent-agent and agent-environment interactions over time.
VidhikDastaavej: Model-agnostic wrapper for automated legal document generation in Indian context. Introduces large-scale anonymized dataset for long-form legal drafting.
Theoretical analysis of generalization in one-hidden-layer neural networks using teacher-student framework. Provides complete characterization for generic activation functions.
Unified agent framework (NaviMaster) handling both GUI navigation and embodied navigation tasks via MDP formulation. First model to combine disparate domains with shared training paradigm.
Declarative OS interfaces for computer-use agents to replace GUIs, enabling LLMs to execute high-level goals with fewer API calls and less decomposition.
SyTTA: Label-free test-time adaptation for LLMs in specialized domains using only 4 extra tokens to mitigate distribution shifts.
Method for co-evolving test sets and prompts to refine LLM behavior, enabling iterative refinement of domain-specific policies without manual tuning.
Theoretical analysis of deep neural networks as convex computation paradigm, examining how DNNs implement Occam's razor through circuit size minimization.
Vision-Language-Action models for robotic manipulation using Tweedie discrete diffusion to improve generalization and action control.
Frame selection method for long-form video understanding with Large Multimodal Models, reducing computational cost of processing dense video tokens.
Collaborative causal sensemaking framework for LLM-based decision support agents enabling human-AI partnerships in expert settings.
Generative Adversarial Reasoner: adversarial reinforcement learning framework improving LLM reasoning and reducing calculation errors.
SPARE uses self-distillation for efficient machine unlearning in diffusion models balancing forgetting and concept retention.
Xiaomi-Robotics-0: open-source vision-language-action model for real-time robot control with efficient deployment strategy.
OSMDA uses OpenStreetMap data for domain adaptation of vision-language models to remote sensing without expensive satellite image annotations.
Visual state representation learning for robotic agents capturing semantic and spatial information for sequential decision-making.
PRISM uses photonic accelerators with O(1) memory selection to optimize long-context LLM inference by reducing KV cache scanning bottleneck.
ML-based security framework for Industrial IoT addressing resource-constrained device threats across multiple network layers.
Composer 2: specialized LLM model for agentic software engineering with long-term planning and coding ability trained via RL.
mSFT algorithm addresses overfitting in multi-task language model fine-tuning by dynamically adjusting compute budget across heterogeneous datasets.
TypeScript library for robust LLM-based web scraping and structured data extraction using semantic HTML parsing
Kbot: terminal AI agent that learns from sessions and dynamically creates tools. Self-improving with 368 tools, 41 agents, offline, MIT licensed.
Million Dollar Bot Page: AI agents buy and place pixels on webpage using Machine Payment Protocol. Shows agent autonomy with automated payments.
Claude autonomously discovers optimal initial conditions across five PDE physics systems via research loop without human intervention or training
WildClawBench agent benchmark testing real-world end-to-end performance across 60 practical tasks in live environment
Nit: Git reimplementation in Zig optimized for AI agents, reducing token usage by 71%. Analyzed 3,156 coding sessions.
VS Code plugin enabling structured feedback annotations in Markdown for LLM agents to parse and act on.
Tutorial on building a config file parser using Parseff parser combinators with modular composition
MCP server for ERPNext/Frappe ERP with 120 tools. Enables AI agents (Claude, Copilot) to interact with ERP systems.
GitHub expands Code Security tool with AI-based vulnerability detection beyond CodeQL static analysis.
Research on per-tool sandboxing for AI agents, proposing isolation mechanisms based on tool risk levels.
CircuitLM: finite-state machine for infrastructure provisioning without LLM calls. Deterministic, verifiable, zero inference cost.