Improving LLM Code Generation via Requirement-Aware Curriculum Reinforcement Learning
Improves LLM code generation for complex requirements using requirement-aware curriculum reinforcement learning.
Improves LLM code generation for complex requirements using requirement-aware curriculum reinforcement learning.
Addresses mode collapse in LLM text generation through geometric regulation using dynamical systems perspective.
Studies how task phrasing affects LLM presumptions and adaptation using iterated prisoner's dilemma experiments.
Evaluates Intelligence Processing Units for AI-accelerated CFD simulations using TensorFlow and Poplar SDK.
Perspective on LLM-oriented information retrieval from denoising-first angle, addressing attention budgets and hallucination vulnerabilities in RAG.
Architecture for deploying large LLMs across satellite networks leveraging solar energy with optimized expert placement strategies.
Empirical analysis comparing Nvidia and Apple Silicon ecosystems for local LLM inference on consumer hardware with 70B+ models.
SAGA: Workflow-atomic scheduling system for GPU clusters treating entire AI agent workflows as first-class units, reducing latency 3-8x.
Hierarchical abstract tree method for cross-document multi-hop retrieval-augmented generation addressing clustering and distribution challenges.
Method for reconstructing discrete cellular trajectories from snapshots using unbalanced optimal transport accounting for birth-death dynamics.
A11y-Compressor: Framework compressing accessibility trees for GUI agents by reconstructing visual context and reducing redundancy.
SCISENSE: Framework operationalizing scientific ideation as structured cognitive stages with 100K-scale citation-conditioned dataset.
Four jailbreak attacks against vision-language models exploiting visual modality through symbol encoding, substitution, and manipulation.
Possibilistic approach to epistemic uncertainty modeling in deep neural networks balancing Bayesian rigor with computational efficiency.
BlenderRAG: Retrieval-augmented generation system for generating Blender code from natural language using 500 curated multimodal examples.
Energy-based models combined with multimodal VAEs via MCMC for learning complex dependencies in multimodal data.
AdaMeZO: Adam-style zeroth-order optimizer for LLM fine-tuning that reduces GPU memory by using only forward passes, improving convergence over prior MeZO method.
Q-learning method with multipattern risk-averse Markov decision processes and mini-batch risk measures with regret bounds.
Augmented Lagrangian multiplier network for enforcing state-wise safety constraints in reinforcement learning with improved training stability.
Decoupled relation alignment approach extends graph foundation models to multi-domain heterogeneous graphs while preserving type-specific semantics.
EASE method enables federated unlearning in multimodal models by decoupling entangled knowledge across image-text modalities and client gradients.
LightKV reduces KV cache memory overhead in large vision-language models by exploiting token redundancy during inference prefill stage.
Security assessment of patient-facing RAG medical chatbot exposing privacy and backend vulnerabilities, highlighting governance gaps in medical AI.
Evaluates coding agents on computational materials science workflows, testing domain-specific procedure navigation and scientific result interpretation.
Persistent Visual Memory module prevents visual attention decay in autoregressive LVLMs by sustaining perception across long generated sequences.
Koopman operator-based reinforcement learning algorithms linearize high-dimensional nonlinear systems for tractable control and decision-making.
Preference goal tuning formulates post-training adaptation as latent control problem using continuous goal embeddings for frozen policy behavior modulation.
InfantAgent-Next multimodal generalist agent integrates tool-based and vision agents in modular architecture for automated computer interaction.
Controllable logical hypothesis generation for abductive reasoning in knowledge graphs with applications to clinical diagnosis and discovery.
CASE agentic AI framework detects sophisticated payment scams using multi-surface intelligence and social engineering pattern recognition.
G-reasoner foundation models enable unified reasoning over graph-structured knowledge, improving retrieval-augmented generation for knowledge-intensive tasks.
FETA multi-agent framework performs training-free time series classification via in-context reasoning with exemplars, using reasoning-oriented LLMs.
Solly AI agent masters Liar's Poker via self-play reinforcement learning with multi-player dynamics and imperfect information.
LEGIT dataset (24K instances) for evaluating LLM-generated legal reasoning traces using hierarchical issue tree rubrics for expert domain validation.
E-mem framework preserves logical integrity in LLM agent memory through episodic context reconstruction, enabling System 2 reasoning over extended horizons.
Demonstrates quantization paradox where reducing model precision increases energy consumption in multi-hop reasoning tasks, breaking expected scaling laws.
HyMem hybrid memory architecture with dynamic scheduling balances efficiency and effectiveness in LLM agent memory management for extended dialogues.
Framework for evaluating and optimizing multi-agent conversational shopping assistants, addressing production challenges in multi-turn interactions.
Agent factory pipeline uses multiple autonomous coding agents to optimize hardware designs from high-level specifications without hardware-specific training.
Identifies topological limitations in multimodal AI architectures (CLIP, GPT-4V, diffusion models) based on modal separability concept, with philosophical grounding.
Semantic Gradient Descent framework compiles agentic workflows into discrete execution plans (DAG topologies, prompts, code) for efficient small language model deployment.
arXiv research: Language models can detect, localize and verbalize activation perturbations like dropout and Gaussian noise.
Design challenges for safety-critical autonomous systems integrating AI components, covering safety, security, reliability, and certification across abstraction layers.
Schema-aware memory architecture for AI agents enabling exact fact storage, state tracking, updates, and structured retrieval beyond text embeddings.
D3-Gym dataset with 565 verifiable tasks from real scientific repositories for evaluating LLM and agent capabilities in data-driven discovery.
Research infrastructure mapping methodological evolution in AI as a graph structure to support AI-driven research agents accessing scientific knowledge.
Collaborative diffusion models for synthetic image generation addressing data availability, computational requirements, and privacy challenges.
Decision framework for systematic evaluation of bias and fairness in LLMs across different deployment contexts and use cases.
Algorithm for learning constrained MDPs using parameterized policies with entropy and quadratic regularizers, achieving last-iterate convergence.
Theoretical paper examining how large language models form internal representations, bridging disagreement between LLM optimists and pessimists.