Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature?
Evaluates whether LLMs are susceptible to biased framing in medical literature abstracts that employ spin techniques.
Evaluates whether LLMs are susceptible to biased framing in medical literature abstracts that employ spin techniques.
PODS decouples rollout generation from policy updates in LLM reinforcement learning, addressing compute and memory asymmetry.
SweRank uses code ranking with LLMs for software issue localization, reducing latency and cost versus multi-step agentic approaches.
IVY-FAKE provides explainable framework and benchmark for detecting AI-generated images and videos with multidimensional annotations.
Theoretical analysis of transformer attention mechanisms, investigating whether universal simulators of attention exist and are learnable.
Addresses fairness issues in federated learning, proposing FeDa4Fair to evaluate bias at client level across multiple sensitive attributes.
Investigates LLM capabilities for medical decision-making, comparing approaches including evidence-based medicine and experimental data.
Automated generation of dynamical system computational models from natural language text using enhanced SysML diagrams and LLMs.
White-Basilisk proposes a hybrid model combining Mamba layers, linear self-attention, and Mixture of Experts for code vulnerability detection.
POLIS framework enables LLMs to accumulate knowledge through multi-agent interaction and inference, mimicking cumulative cultural evolution.
FLOSS framework enables federated learning with user opt-out and stragglers support, addressing data privacy in heterogeneous distributed systems.
ReasonRank method empowering passage ranking with reasoning ability by leveraging large reasoning models for improved listwise ranking in complex scenarios.
Establishes task-stratified knowledge scaling laws for post-training quantized LLMs, analyzing quantization impact on memorization, application, and reasoning capabilities.
SMARTER framework for explainable toxicity detection using LLMs with synthetic explanation generation and preference optimization for content moderation.
Studies transformer capability to learn transitive relation reasoning in graphs, essential for LLM factual correctness and causal inference tasks.
Multi-Level Optimal Transport (MOT) framework for aligning representational structures across model layers and brain regions with global alignment scoring.
Locate-Then-Examine two-stage VLM-based forensic framework for detecting AI-generated images using grounded region reasoning for artifact identification.
Reframes human label variation in NLP from noise to signal for improving model robustness, particularly relevant for post-training methods with human feedback.
Pilot study testing hypothesis that continual exposure to junk web text induces cognitive decline in LLMs using controlled Twitter/X corpus experiments.
CodeRL+ improves LLM code generation using reinforcement learning with execution semantics alignment, bridging gap between text patterns and functional correctness.
Analysis showing LLMs develop universal sinusoidal representations of numbers across different families, with representations largely interchangeable across models.
OpenHands Software Agent SDK provides a composable, extensible toolkit for building production-ready software engineering agents with flexible implementation and secure execution.
ItemRAG applies retrieval-augmented generation with item-based similarity to improve LLM-based recommendation systems, addressing cold-start problems.
AutoGraphAD uses variational graph autoencoders for unsupervised network intrusion detection without requiring labeled datasets.
Digital in-memory stochastic computing architecture using compressed bent-pyramid format to optimize AI model matrix multiplication operations.
Hybrid-AIRL method combining adversarial inverse reinforcement learning with supervised expert guidance, evaluated on poker game with imperfect information.
Research on sparse dictionary learning in mechanistic interpretability, analyzing how neural networks represent concepts as linear directions and encode multiple concepts.
Device-native autonomous agent system for privacy-preserving automated negotiations in insurance and B2B commerce without centralized servers.
CEDAR agentic system automating data science tasks via context engineering, handling complexity, data size, and computational constraints.
SciCoQA dataset of 635 paper-code discrepancies to evaluate LLM capability for cross-modal verification and research reproducibility auditing.
KOCO-BENCH benchmark evaluating how LLMs acquire and apply domain knowledge in specialized software development tasks.
Theoretical work relaxing realizability assumptions in language identification and generation tasks, establishing statistical rates without distribution constraints.
QuantaAlpha is an evolutionary LLM-driven agentic framework for financial alpha mining with multi-round search and experience reuse capabilities.
NeuroSymActive combines neural-symbolic reasoning with knowledge graphs for complex multi-hop question answering, integrating LLMs with structured knowledge.
CAST system for stable LLM-based text analysis of tabular data with algorithmic prompting to ensure output consistency for data analytics tasks.
Domain-specific LLM application for residential energy retrofit decision-making, guiding homeowners through building energy assessments.
Developer tool for flexible and high-performance fully sharded data parallel training, enabling block-wise quantization and structure-aware methods at scale.
Medical imaging application of vision-language models with pretraining for differential VQA tasks requiring fine-grained visual comparison.
Study analyzing semantic representation alignment between language models and vision encoders in vision-language models for taxonomic generalization.
Research on membership inference attacks against contrastive pretraining models like CLIP to audit PII memorization in multimodal backbones.
AutoML approach combining automated machine learning with deep unfolding for wireless beamforming optimization using learned proximal gradient descent layers.
Research evaluates robustness of climate foundation models under distribution shift, testing generalization beyond training data.
Explainable AI analysis revealing why AI-generated text detectors fail despite high benchmark accuracy, exploiting dataset-specific artifacts.
Visual Masked Autoencoder with Normalizing Flow for time series anomaly detection using foundation models for improved generalization.
CARLA-Air: Unified simulation infrastructure for air-ground embodied AI combining drone and vehicle dynamics in single environment.
Causal analysis showing LLMs fail reasoning when surface heuristics conflict with implicit constraints, studied through car wash problem.
Theoretical analysis of generalization bounds for overparameterized shallow neural networks related to distance from initialization.
Investigation of whether smaller language models with task-aware retrieval can match larger proprietary models for scientific knowledge discovery applications.
QuanBench+: Unified benchmark for evaluating LLMs on quantum code generation across Qiskit, PennyLane, and Cirq frameworks with 42 aligned tasks.
PromptEcho: Reward construction method using vision-language models for text-to-image model RL without human annotations or additional training.