Generalized Leverage Score for Scalable Assessment of Privacy Vulnerability
Generalized leverage score method for assessing privacy vulnerability and membership inference attack risk without model retraining.
Generalized leverage score method for assessing privacy vulnerability and membership inference attack risk without model retraining.
Study of Model Context Protocol design choices for tool orchestration and code execution in multi-server agent systems.
DocSplit introduces first benchmark dataset for document packet splitting task in heterogeneous multi-page document understanding.
Scoping review of 59 publications examining motivations and adoption barriers for dataset documentation tools in AI development.
B-DENSE proposes branching method for dense ensemble learning in diffusion models to improve inference speed and preserve structure.
ReLoop addresses silent failures in LLM-based optimization code generation using structured modeling and behavioral verification.
Study tracks capability emergence mechanisms across neural network scales using five geometric measures on 120+ emergence events.
MAEB benchmark evaluates 50+ audio embedding models across 30 tasks in speech, music, and cross-modal audio-text reasoning.
MedProbCLIP applies probabilistic vision-language models for chest X-ray and radiology report retrieval in biomedical applications.
RCT study (N=979) evaluates prompting instruction interventions teaching students to use GenAI as tutor for better learning outcomes.
ML-CARE proposes a carbon-aware evaluation metric for AI models to measure environmental cost alongside standard performance benchmarks.
Theoretical analysis of how generative AI systems perform when trained on contaminated data containing AI-generated content mixed with human-generated content.
Study of theory-of-mind reasoning in 41 open-weight language models, testing false belief tasks to understand mental state reasoning capabilities across diverse models.
Method for updating LLM knowledge from new documents while retaining post-training capabilities like instruction-following and reasoning through context distillation.
Federated learning approach for detecting cross-border insider threats in government financial systems using graph-based reasoning on distributed, privacy-sensitive data.
Surrogate-based measurement technique using LLM labeling for cost-effective prevalence estimation in large-scale A/B testing.
Learnable index approach for approximate nearest neighbor search in large-scale recommendation systems with joint embedding and indexing.
Analysis of ecosystem-level failure in retrieval systems when AI-generated content pollutes search results and RAG training data.
Empirical study examining how user domain knowledge and AI literacy affect interaction with LLM-integrated building energy management systems.
Multi-listener reinforcement learning approach to improve faithfulness of chain-of-thought reasoning in LLMs while maintaining performance.
Hierarchical reinforcement learning framework with explicit credit assignment for LLM agents on long-horizon multi-turn tasks.
Theoretical framework using convex conjugate duality to analyze trainability and generalization in deep neural networks with SGD.
Training-free approach to adapt language models by identifying and leveraging local module activations without retraining.
Taxonomy and analysis of how LLMs handle long-tail knowledge from infrequent, domain-specific, cultural, and temporal domains.
Graph Neural Network approach for sea ice modeling using collision physics, positioning ice pieces as nodes and interactions as edges.
Study evaluating LLMs as zero-shot annotators for Bangla hate speech detection, examining reliability and bias in low-resource language settings.
Qualitative study of generative AI usage among part-time university students in education and professional contexts.
Analysis of electromagnetic fault injection attacks on embedded neural network models across different number representations.
Graph meta-network architecture for weight-space models that predict neural network accuracy on new datasets.
Self-supervised learning approach for improving feature representations in object detection without labeled data.
HAWX hardware-aware framework for DNN approximation using multi-level sensitivity scoring to guide AxC block integration across abstraction levels.
Framework combining pathology foundation model with transformer decoder for automated histopathology report generation from whole slide images.
Multilingual OCR system for India using vision-language models through Chitrapathak series, balancing linguistic diversity and deployment constraints.
Studies bias spillover effect in LLM fairness alignment, showing single-attribute mitigation can exacerbate disparities in untargeted dimensions.
RoboGene uses agentic framework for automated task generation and curation to improve vision-language-action model pre-training data diversity.
IndicEval benchmarks LLMs on authentic UPSC, JEE, NEET exam questions in English and Hindi across STEM and humanities.
Team-of-Thoughts orchestrates heterogeneous agent models via tool calling to optimize multi-agent system performance through complementary capabilities.
Proposes social meta-learning approach enabling LLMs to proactively solicit and learn from conversational corrective feedback.
Unifies looping and depth-growing in LLMs mechanistically, showing both exhibit convergent depth-wise signatures and improved reasoning.
CALMs balance interpretability of generalized additive models with accuracy by adding conditional pairwise interactions.
RLM-JB framework uses recursive language models for jailbreak detection in agentic systems executing tools over untrusted content.
MerLean automates formalization of quantum computation papers from LaTeX to Lean 4 code, producing 2,050 declarations from 114 statements.
LSTM-based deterministic model for global streamflow forecasting pre-trained on ERA5-Land reanalysis data and fine-tuned on operational forecasts.
DataJoint 2.0 introduces relational workflow model for scientific data pipelines with provenance tracking and transactional guarantees, enabling agentic scientific workflows.
FlowPrefill decouples preemption from prefill scheduling to reduce head-of-line blocking in LLM serving systems. Improves time-to-first-token latency.
Layer-wise integrated gradients method for explaining Transformer predictions with context-awareness. Improves interpretability of deep neural models.
Proposes using LLMs as evaluators for comparative assessment with improved reliability modeling. Addresses bias and inconsistency in LLM judge performance.
Formalizes causal abstraction between low and high-level models using natural transformations. Addresses interpretability and robustness in AI systems.
Systematic evaluation of tokenization strategies for MEG neuroimaging foundation models. Compares discretization approaches for continuous neural time series data.
Convergence analysis for differential temporal difference learning in average reward reinforcement learning settings.