DynaTab: Dynamic Feature Ordering as Neural Rewiring for High-Dimensional Tabular Data
DynaTab: dynamic feature ordering architecture for high-dimensional tabular data using neural rewiring with permutation complexity detection.
DynaTab: dynamic feature ordering architecture for high-dimensional tabular data using neural rewiring with permutation complexity detection.
Analysis of LLM safety bypasses using mathematical encoding (set theory, logic, quantum mechanics) achieving 46-56% attack success across models.
AEMG: self-supervised pre-training framework for learning generalizable action representations from electromyography across subjects and devices.
FINER-SQL boosts small language models for text-to-SQL tasks, enabling efficient on-premise deployment with improved reasoning and instruction following.
Method for detecting sycophancy in conversational AI therapists using dynamic emotional signature graphs without LLM judges.
CuraView multi-agent framework using GraphRAG for detecting hallucinations in LLM-generated medical discharge summaries from EHRs.
MEMSAD: security framework for detecting gradient-based anomalies and memory poisoning attacks on retrieval-augmented LLM agents via Stackelberg game analysis.
MHPR benchmark for evaluating vision-language models on human perception and reasoning tasks across individual, multi-person, and interaction dimensions.
Systematic evaluation challenging whether graph-tokenizing LLMs effectively understand graph structure or merely process tokens.
Modular AI agent pipeline using discrete skills for automating Library of Congress subject indexing in library cataloging.
Benchmark testing language model agents' ability to build complete software projects from scratch with minimal human oversight.
Benchmark for evaluating machine unlearning in vision-language models to remove memorized copyrighted visual content.
KV-cache quantization method measuring error in model-visible coordinates for efficient LLM inference with low-rank residuals.
Empirical study of LLMs as agents in repeated games, testing strategic behavior from international relations theory.
ELAS enables efficient LLM pre-training via low-rank adaptation combined with 2:4 activation sparsity for GPU acceleration.
Open-vocabulary semantic mapping for robots using 3D fusion of voxel and instance-level semantic embeddings.
LLM-based framework for smart contract vulnerability detection with vulnerability-specific prompts and 31K dataset.
SERE method uses structural example retrieval to enhance LLM performance on event causality identification and reduce causal hallucination.
SAM-NER framework for zero-shot NER using semantic archetype mediation to handle domain and schema shifts in LLMs.
Algorithms for segmenting human-written and LLM-generated text in co-authored documents using change point detection.
Analysis of LoRA rank threshold requirements for fine-tuning, examining theoretical conditions and practical implications for cross-entropy loss.
Workflow framework for asynchronous human-AI collaboration in HPC environments with checkpoint-based intervention for high-stakes applications.
Identifies under-memorization failures in LVLM unlearning benchmarks that use fictitious identities for privacy evaluation.
Case study on upskilling software teams with AI Advocates to enable human-AI collaborative development.
RoboAlign-R1 uses reward alignment to improve robot video world models for instruction following and manipulation tasks.
TRACE framework for trustworthy agentic AI systems in critical domains, combining architecture, metrics, and human supervision protocols.
MCJudgeBench benchmark evaluates LLM judges on constraint-level accuracy in multi-constraint instruction following tasks.
Train-free dataset distillation using semantic-distribution matching in diffusion models without fine-tuning.
Deco dual-embodiment framework extending emotional bonds from physical objects to AI companion agents.
Framework to improve activation steering in LLMs by distilling prompt-based steering behavior into interpretable models.
Randomized trial showing atomic fact-checking increases clinician trust in LLM oncology decision support recommendations.
Language models perform iterative conceptual analysis through counterexample generation and definition repair chains.
iWorld-Bench comprehensive benchmark for training and evaluating interactive world models with unified action generation framework.
TabSurv adapts modern tabular neural networks to survival analysis using Weibull and non-parametric methods.
MOSAIC-Bench benchmark measuring safety vulnerabilities in coding agents through compositional task decomposition.
MAKA: Multi-agent human-in-the-loop system for CNC aerospace manufacturing combining LLM text generation with risk-constrained numerical workflows and auditable provenance.
SaFE-Scale framework showing clinical LLM safety and accuracy follow different scaling laws; safety doesn't necessarily improve with model size or context.
Meta-review database and taxonomy of AI risks including terminology standardization across research, policymakers, and companies for shared risk discussion.
Position paper on safety requirements for open-ended AI agents with autonomous behavior generation and self-evolution capabilities before deployment.
COMPASS: Multi-agent framework integrating vision-language models for closed-loop planning in cooperative multi-agent RL, improving sample efficiency and generalization.
Agentic Publications: LLM-driven framework transforming scientific papers into interactive knowledge systems integrating structured data and multimedia.
Position paper arguing tool-augmented agents should invoke external tools only when epistemically necessary, not treating tools as ordinary actions.
VCBench: First benchmark for predicting founder success in venture capital using LLMs, with sparse signals and uncertain outcomes exceeding market index performance.
Combines LLM agents with operations research algorithms for inventory control, enabling flexible reasoning while adapting to demand distribution shifts.
HiMAC: Hierarchical macro-micro learning framework enabling LLM agents to handle long-horizon tasks with structured planning and reliable execution via separate reasoning/action generation.
Framework for quantifying trust in autonomous AI agents through end-to-end operational outcomes rather than model-internal properties, relevant to deployed agents with financial risk.
Soft Tournament Equilibrium: Framework for evaluating non-transitive LLM-based agents using set-valued cores instead of linear rankings for cyclic competitive domains.
HiL-Bench evaluates whether coding agents know when to ask for help versus act autonomously on incomplete specifications, addressing judgment gaps in frontier agents.
HWE-Bench: First large-scale benchmark with 417 real hardware bug repair tasks for evaluating LLM agents at repository-level, beyond component-level HDL generation.
Mechanized proofs in Coq for structural governance of cognitive workflows using coinductive logic and the Interaction Trees library.