Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning
Odysseus extends vision-language models to 100+ turn decision-making in games using reinforcement learning, improving long-horizon performance.
Odysseus extends vision-language models to 100+ turn decision-making in games using reinforcement learning, improving long-horizon performance.
MemRouter decouples memory management from LLM generation in conversational agents using embedding-based routing for long-term memory decisions.
Text mining analysis of ChatGPT research publications in programming education, identifying four dominant themes in scholarly discourse.
AlphaInventory applies LLM-based evolutionary search to optimize inventory policies in online, non-stationary environments with deployment guarantees.
Large multimodal model for music understanding combining audio encoders with mixture-of-experts design for time-series and non-time-series music tasks.
Benchmark study of social bias in LLM-generated code across 343 real-world tasks, extending prior Solar work with SocialBias-Bench evaluation framework.
Agent Capsules runtime optimizes multi-agent LLM pipelines by merging agents adaptively while maintaining quality.
RadLite uses LoRA fine-tuning on small language models for radiology tasks deployable on consumer CPUs.
BWLA method achieves 1-bit weight and activation quantization for LLMs, enabling efficient deployment.
Proposes trust schema and verification framework for agent skills as deployable artifacts in LLM agent runtimes.
Improves LLM code generation for complex requirements using requirement-aware curriculum reinforcement learning.
Addresses mode collapse in LLM text generation through geometric regulation using dynamical systems perspective.
Studies how task phrasing affects LLM presumptions and adaptation using iterated prisoner's dilemma experiments.
Evaluates Intelligence Processing Units for AI-accelerated CFD simulations using TensorFlow and Poplar SDK.
Perspective on LLM-oriented information retrieval from denoising-first angle, addressing attention budgets and hallucination vulnerabilities in RAG.
Architecture for deploying large LLMs across satellite networks leveraging solar energy with optimized expert placement strategies.
Empirical analysis comparing Nvidia and Apple Silicon ecosystems for local LLM inference on consumer hardware with 70B+ models.
SAGA: Workflow-atomic scheduling system for GPU clusters treating entire AI agent workflows as first-class units, reducing latency 3-8x.
Hierarchical abstract tree method for cross-document multi-hop retrieval-augmented generation addressing clustering and distribution challenges.
Method for reconstructing discrete cellular trajectories from snapshots using unbalanced optimal transport accounting for birth-death dynamics.
A11y-Compressor: Framework compressing accessibility trees for GUI agents by reconstructing visual context and reducing redundancy.
SCISENSE: Framework operationalizing scientific ideation as structured cognitive stages with 100K-scale citation-conditioned dataset.
Four jailbreak attacks against vision-language models exploiting visual modality through symbol encoding, substitution, and manipulation.
Possibilistic approach to epistemic uncertainty modeling in deep neural networks balancing Bayesian rigor with computational efficiency.
BlenderRAG: Retrieval-augmented generation system for generating Blender code from natural language using 500 curated multimodal examples.
Energy-based models combined with multimodal VAEs via MCMC for learning complex dependencies in multimodal data.
AdaMeZO: Adam-style zeroth-order optimizer for LLM fine-tuning that reduces GPU memory by using only forward passes, improving convergence over prior MeZO method.
Q-learning method with multipattern risk-averse Markov decision processes and mini-batch risk measures with regret bounds.
Augmented Lagrangian multiplier network for enforcing state-wise safety constraints in reinforcement learning with improved training stability.
Decoupled relation alignment approach extends graph foundation models to multi-domain heterogeneous graphs while preserving type-specific semantics.
EASE method enables federated unlearning in multimodal models by decoupling entangled knowledge across image-text modalities and client gradients.
LightKV reduces KV cache memory overhead in large vision-language models by exploiting token redundancy during inference prefill stage.
Security assessment of patient-facing RAG medical chatbot exposing privacy and backend vulnerabilities, highlighting governance gaps in medical AI.
Evaluates coding agents on computational materials science workflows, testing domain-specific procedure navigation and scientific result interpretation.
Persistent Visual Memory module prevents visual attention decay in autoregressive LVLMs by sustaining perception across long generated sequences.
Koopman operator-based reinforcement learning algorithms linearize high-dimensional nonlinear systems for tractable control and decision-making.
Preference goal tuning formulates post-training adaptation as latent control problem using continuous goal embeddings for frozen policy behavior modulation.
InfantAgent-Next multimodal generalist agent integrates tool-based and vision agents in modular architecture for automated computer interaction.
Controllable logical hypothesis generation for abductive reasoning in knowledge graphs with applications to clinical diagnosis and discovery.
CASE agentic AI framework detects sophisticated payment scams using multi-surface intelligence and social engineering pattern recognition.
G-reasoner foundation models enable unified reasoning over graph-structured knowledge, improving retrieval-augmented generation for knowledge-intensive tasks.
FETA multi-agent framework performs training-free time series classification via in-context reasoning with exemplars, using reasoning-oriented LLMs.
Solly AI agent masters Liar's Poker via self-play reinforcement learning with multi-player dynamics and imperfect information.
LEGIT dataset (24K instances) for evaluating LLM-generated legal reasoning traces using hierarchical issue tree rubrics for expert domain validation.
E-mem framework preserves logical integrity in LLM agent memory through episodic context reconstruction, enabling System 2 reasoning over extended horizons.
Demonstrates quantization paradox where reducing model precision increases energy consumption in multi-hop reasoning tasks, breaking expected scaling laws.
HyMem hybrid memory architecture with dynamic scheduling balances efficiency and effectiveness in LLM agent memory management for extended dialogues.
Framework for evaluating and optimizing multi-agent conversational shopping assistants, addressing production challenges in multi-turn interactions.
Agent factory pipeline uses multiple autonomous coding agents to optimize hardware designs from high-level specifications without hardware-specific training.
Identifies topological limitations in multimodal AI architectures (CLIP, GPT-4V, diffusion models) based on modal separability concept, with philosophical grounding.