TravelFraudBench configurable framework for evaluating GNNs on fraud ring detection across travel-specific topologies with mechanisms for diverse fraud patterns.
Framework using fine-tuned multimodal LLM (Gemma 3 27B) to assess building conditions from street-view imagery, outperforming human raters.
AGNT2 designs blockchain Layer 2 infrastructure optimized for high-frequency autonomous agent interactions with identity, escrow, and state management.
CSTM-Bench dataset and algorithms for cross-session threat detection in AI agents, addressing memoryless guardrails vulnerable to distributed attacks.
Multi-task learning system for automated discourse analysis classifying teacher and student utterances along reasoning and utterance-type dimensions.
Proposes machine mental imagery representations for maintaining shared context in situated dialogue beyond immediate context windows for conversational agents.
Quantifies LLM demographic bias from implicit linguistic signals versus explicit user profiles, examining how identity is conveyed through socio-linguistic factors.
Adaptive instruction composition method for LLM red-teaming that combines crowdsourced tactics with attacker LLM to discover diverse jailbreaks more effectively.
Systematic investigation of Gaussian scale parameter effects in Gaussian Kolmogorov-Arnold Networks through first-layer geometry and approximation behavior.
Combines SAT solving with LLM-generated code to discover infinite families of Ramsey-good graphs, demonstrating computer-assisted mathematical discovery.
Analysis of vision-language-action models in robotics for long-horizon tasks, examining how success metrics evaluate complex real-world manipulation problems.
Novel two-step reasoning approach using LLMs to enhance occupation prediction by generating reasons from career history before making recommendations.
SQLyzr is a benchmark platform for evaluating text-to-SQL models built with LLMs, providing granular evaluation across query types beyond single aggregate scores.
IRM proposes zero-shot LLM-generated text detection via implicit reward models without training data.
EngramaBench benchmarks long-term conversational memory in LLMs using graph retrieval across multi-session conversations.
SparKV adaptive KV cache loading framework combining cloud-based streaming with on-device computation for efficient LLM inference on resource-constrained devices.
CorridorVLA predicts sparse spatial anchors as explicit constraints for vision-language-action models to guide action generation with tolerance regions.
CAP method for selective knowledge unlearning in LLMs via controllable alignment prompting without modifying model weights, enabling closed-source model compliance.
Calibrated prediction-powered inference method for semisupervised mean estimation with miscalibrated black-box predictors, extending augmented inverse-probability weighting.
Co-evolving proposer and visual critic framework using reinforcement learning for precise GUI grounding, mapping natural language to pixel coordinates.
Benchmarks fairness of LLM-based speech recognition decoders compared to task-specific and implicit language models across demographic groups.
Investigates synthetic data augmentation for controllable human-centric video generation to address dataset scarcity for rare identities and complex actions.
MiMIC framework addresses visual modality collapse in universal multimodal retrieval while preventing semantic misalignment between vision and text embeddings.
Analysis of test-time reinforcement learning for math reasoning showing spurious signals from label noise are amplified by group-relative advantage, proposes mitigation strategies.
PolyChartQA dataset with 534 multi-chart images for question-answering tasks requiring interpretation of related charts together.
Self-supervised learning method for aerial imagery using additive-residual selective invariance to handle degraded images (haze, blur, rain) without enforcing spurious alignment.
Fine-tuned LLM approach for detecting machine-generated code snippets across programming languages, including source attribution and hybrid human-machine code detection.
VLAA-GUI framework for autonomous GUI agents with modules for stopping verification, action recovery, and search to prevent premature success and repetitive loops.
Interactive retrieval-augmented approach (IRAP) for converting natural language software performance requirements into mathematical specifications using preference elicitation.
Mathematical analysis proving supervised learning has geometric constraints that limit encoder representation, causing necessary sensitivity to label-correlated but test-time irrelevant features.
VG-CoT framework for trustworthy visual reasoning in large vision-language models through grounded multi-step chain-of-thought with explicit region alignment.
CSC defense method against backdoor poisoning attacks in DNNs that leverages adversary's poison data to improve robustness without accuracy degradation.
Evaluates NER-based automated de-identification of clinical notes with differential privacy guarantees for GDPR/HIPAA compliance.
VARestorer applies one-step VAR distillation to visual autoregressive models for real-world image super-resolution by improving global context exploitation.
Studies reasoning primitives (recall, state-tracking) in hybrid vs attention-only LLM architectures, evaluating suitability for joint reasoning tasks.
DP-RL framework augments policy gradient learning with dynamical prior loss to improve temporal coherence and prevent degenerate behaviors in RL policies.
MISTY high-throughput motion planner using mixer architectures for single-step trajectory generation in autonomous driving with low latency.
TaNOS framework for robust numerical reasoning over tables through header anonymization and operation sketches to reduce memorization and improve domain generalization.
Attention-based multiple instance learning framework with foundation models for lung adenocarcinoma growth pattern prediction from pathology images.
Knowledge distillation approach combining pre-trained LLMs with sequential recommenders for efficient user understanding and real-time inference.
HAF-DS framework integrates LSTM forecasting with optimization for coupled demand-supply prediction in volatile supply chains.
Metamorphic testing approach to detect memorization and data leakage in LLM-based automated program repair systems.
Investigates role of public test cases in multi-agent LLM-driven code generation frameworks for algorithmic problem-solving and debugging.
Studies preprocessing and memristor dynamics in reservoir computing for image classification, exploring hardware efficiency of recurrent architectures.
Verbal Process Supervision framework guides LLM reasoning through structured natural-language critique in iterative generate-refine loops, improving performance on mathematical benchmarks.
Compares n-gram models with LSTM/Transformer architectures for event-log prediction, finding lightweight automata achieve comparable accuracy with fewer resources.
Multi-task RL for autonomous underwater vehicle control using subnetwork discovery to enable adaptive, interpretable policies under uncertain conditions.
Backdoor attack methodology against LLMs using natural style triggers with reliable payload injection and threat model specification.
Open datasets and benchmarks for precise video captioning using structured visual specification and human-AI oversight.
Agent Evolving Learning: framework enabling LLM agents to learn from past episodes and improve behavior in open-ended environments.