Chain of Mindset: Reasoning with Adaptive Cognitive Modes
Chain of Mindset method enabling LLMs to adaptively switch between cognitive modes for improved reasoning across problem-solving stages.
Chain of Mindset method enabling LLMs to adaptively switch between cognitive modes for improved reasoning across problem-solving stages.
HECG framework for autonomous agents using LLM-based action generation with multi-dimensional strategy alignment and error correction mechanisms.
JobMatchAI production system for job candidate matching using Transformer embeddings, skill knowledge graphs, and explainable AI reranking.
Comprehension-Gated Agent Economy architecture gates AI agent economic agency on verified comprehension functions rather than capability benchmarks.
Latent Posterior Factors framework for aggregating noisy heterogeneous evidence in decision-making with explicit uncertainty quantification and interpretability.
Theoretical characterization of Latent Posterior Factors framework for aggregating multiple evidence sources in probabilistic reasoning with formal guarantees.
Model Workspace Protocol for sequential agentic workflows using folder structure as architecture, reducing engineering overhead versus multi-agent frameworks.
TRUST-SQL uses multi-turn reinforcement learning and tool integration for text-to-SQL parsing on unknown database schemas in enterprise settings.
Theoretical analysis of delta-margin majority voting for consensus-based prediction quality control in high-stakes ML applications.
Uses reinforcement learning with learned gadgets to design quantum circuits addressing noise and connectivity constraints in real quantum hardware.
ACT-JEPA combines joint-embedding predictive architecture with self-supervised learning for efficient policy representation in imitation learning.
Physics-Informed Evolution framework embeds physical laws into evolutionary algorithm fitness functions for quantum control problems.
Oracular Programming framework for building modular, composable LLM-enabled software with enforceable contracts and reliable composition primitives.
Method to improve mathematical reasoning in smaller LLMs by integrating arithmetic learning alongside knowledge distillation and data augmentation.
Proposes frequency progressive autoregressive approach for image generation using continuous tokens instead of raster-scan prediction.
Research on minimal data repair showing imputing all missing values unnecessary for accurate ML models; introduces minimal and almost-minimal repair concepts.
SocialJax is an evaluation suite for multi-agent reinforcement learning in sequential social dilemmas, measuring generalization across social scenarios.
Study of data deduplication effects on deep neural network image classifiers. Examines impact on robustness against adversarial attacks.
Systematic review of uncertainty quantification and mitigation methods for LLMs to address hallucination and calibration challenges.
Survey of edge-cloud collaborative computing for AI and LLM deployment. Covers distributed intelligence and model optimization techniques.
RAGXplain framework for evaluating and debugging RAG systems. Provides actionable insights into retrieval, context, and generation performance beyond aggregate metrics.
Survey of LLM-based software quality assurance techniques covering requirement analysis, code review, test generation, and standards compliance.
Benchmark for text-to-SQL systems evaluating scientific reasoning over biomedical knowledge bases requiring implicit domain understanding.
Analysis revealing association biases beyond token level cause LLM content moderation over-sensitivity on decontextualized statements.
Offline RL method handling dynamics mismatch between source and target datasets by leveraging high-shift regions for better exploration.
Self-organizing maps extension addressing catastrophic forgetting in continual learning through saturation mechanisms.
Benchmark evaluating memory capabilities in LLM agents including memorization, updating, and retrieval of long-term information across multi-turn interactions.
Protocol-agnostic tool management library for function-calling LLMs that reduces fragmentation and development overhead in LLM applications.
Method for predicting better pre-trained weights via retrodiction of forgetting to encapsulate more knowledge beyond training datasets.
Analysis of fast weight programmers and 2D-state RNNs as linear transformers with connections to neurobiology and language modeling advances.
Framework for optimizing content for generative search engines powered by LLMs and RAG by modeling user intent and role-based search patterns.
Study showing psychometric questionnaires designed for humans mischaracterize LLM psychology compared to behavior in real user interactions.
Object-centric evaluation framework for automated fine-grained assessment of multi-turn instruction-based image editing using VLMs.
Tree-based group relative policy optimization for LLM agent reinforcement learning addressing sparse supervision in long-horizon multi-turn tasks.
Method aligning supervised fine-tuning with in-context learning activations to improve LLM generalization and calibration in data-scarce settings.
Offline RL framework using linear Transformers for compositional Q-function estimation across diverse subtasks via in-context learning.
Parameter-efficient machine unlearning method for foundation models addressing weight unbounding in privacy removal.
Benchmark evaluating metrics and judges for assessing harmful content generation in LLMs.
Slow-Fast Policy Optimization framework improving LLM reasoning via RL with better gradient stability and exploration.
Methods for detecting data contamination in LLM evaluation during reinforcement learning post-training phases.
Scalable energy-based models via adversarial training unifying classification and generative modeling using EBM framework.
CBF-RL integrates control barrier functions into reinforcement learning training for enforcing safety constraints during policy training.
Frame semantic patterns methodology for identifying underreported gender-based violence in e-medical records using NLP.
Generative hints training methodology enforcing functional invariances in vision models beyond empirical training data.
Neural network approach to accelerate computation of n-particle reduced density matrices for strongly-correlated quantum states.
Methodology for sustainable ML model evaluation addressing gaps in Green AI auditing practices and standardization.
Research showing genomic sequence models exhibit in-context learning similar to LLMs, demonstrating ICL emerges across sequence domains.
Uni-DAD: Unified method for distilling and adapting diffusion models for few-step image generation in new domains.
Benchmark for evaluating multimodal LLMs on schema-grounded visual information extraction and reasoning tasks in agentic settings.
Framework for runtime monitoring of multi-agent systems using cryptographic provenance and drift detection to mitigate emergent norms at scale.