Concept Inconsistency in Dermoscopic Concept Bottleneck Models: A Rough-Set Analysis of the Derm7pt Dataset
Analyzes concept inconsistencies in dermoscopic concept bottleneck models using rough set theory on medical imaging dataset.
Analyzes concept inconsistencies in dermoscopic concept bottleneck models using rough set theory on medical imaging dataset.
Empirical study examining active learning for chemical reaction extraction from literature, identifying annotation cost barriers.
Extends online federated learning with stochastic adversary model capturing statistical variation and enabling parallelization benefits.
Formulates language model scaling for scientific discovery around evaluation-driven loops with verifiers and task-specific scoring functions.
Proposes LASER framework formulating active sensing for field reconstruction as POMDP with closed-loop adaptive sensor placement.
Introduces FairTree algorithm for auditing machine learning models for subgroup fairness using bias-variance decomposition framework.
Proposes TACENR for task-agnostic explainability of node representations in graph neural networks using contrastive methods.
Studies optimal routing for federated learning over dynamic satellite networks with relay-based communication between clients and servers.
Analyzes catastrophic forgetting in continual knowledge graph embedding updates and proposes approaches beyond limiting embedding changes.
Introduces unsupervised method for calibrating confidence estimates in reasoning LLMs using single generation without labels or repeated sampling.
Proposes personalized federated learning approach for industrial predictive analytics accounting for heterogeneous degradation processes across clients.
Proposes EVPO method for LLM post-training via reinforcement learning that optimizes critic utilization in policy gradient optimization.
Evaluates GNN performance on Bitcoin fraud detection under strict inductive protocols, finding random forests outperform GCN, GraphSAGE, and GAT methods.
Reviews decentralized optimization approaches for distributed machine learning across devices with local datasets, emphasizing privacy and scalability benefits.
Comparative analysis of RaBitQ and TurboQuant quantization methods with reproducible experimental framework showing inconsistent performance differences.
Proposes stochastic attention modification for transformer-based scientific models to improve calibrated uncertainty estimates at inference time.
Theoretical analysis separating geometry from probability in machine learning generalization bounds for in-sample and out-of-sample performance.
Framework for structure-based drug discovery combining contrastive 3D protein-ligand learning with autoregressive molecular generation on commercial compounds.
Theoretical analysis of constant-stepsize Q-learning using stochastic switching systems and Lyapunov methods for convergence certification.
Black-box reduction from online learning to multicalibration using no-regret learners and expected variational inequality solvers.
SAGE method for edge-cloud hybrid inference under hard uplink bandwidth constraints using semantic evidence composition instead of attention-based importance.
Self-supervised representation learning framework for structural damage identification that disentangles damage from operational variability without labels.
HardNet++ enforces nonlinear constraints in neural network outputs for control and decision-making with guaranteed constraint satisfaction at inference.
PREF-XAI framework generates personalized rule-based explanations for black-box ML models based on user preferences and cognitive constraints.
Adaptive MSD-Splitting algorithm improves discretization of continuous attributes in decision trees like C4.5 and Random Forests for efficiency and accuracy.
Theoretical analysis of adversarial training in Vision Transformers, studying robustness to adversarial examples and benign overfitting.
FB-NLL: Feature-based approach handling noisy labels in personalized federated learning with heterogeneous data.
FASTER: Value-guided sampling method reducing computational cost of test-time scaling in diffusion-based RL policies.
Safe continual RL for non-stationary environments ensuring constraint satisfaction during learning in physical control systems.
Method for scaling test-time compute in agentic coding systems using trajectory ranking and value estimates for long-horizon tasks.
Dual Triangle Attention: Bidirectional transformer attention mechanism eliminating need for positional embeddings.
Agent-GWO: Multi-agent collaborative system for dynamic prompt optimization in LLMs using genetic wolf optimizer.
Analysis showing indistinguishability properties don't prevent data extraction from LLM APIs; formalizes privacy game separation.
Empirical study of jailbreak detection in LLMs using multi-generation sampling on JailbreakBench with varying model alignment.
ARES: Red-teaming and repair method for RLHF systems addressing vulnerabilities in both policy and reward model simultaneously.
Evaluation of LLM-based scientific agents across 8 domains with 25k runs assessing whether they follow scientific reasoning norms.
Framework testing LLM sensitivity to semantic changes in document comparison tasks using perturbation-based needle-in-haystack experiments.
Hierarchically robust zero-shot vision-language models defending against adversarial attacks on both base and superclass levels.
Empirical study of prompt engineering for formal mathematical reasoning, analyzing 40+ prompt variants in equational theory competition.
MORPHOGEN: multilingual benchmark for evaluating LLM handling of gender-aware morphological generation in 13 languages.
Framework for personalized LLM benchmarking that evaluates alignment with individual user preferences instead of aggregate ratings.
R²-dLLM: acceleration technique for diffusion-based LLMs via spatio-temporal redundancy reduction during parallel token decoding.
TRN-R1-Zero: zero-shot reasoning on text-rich networks using LLMs with reinforcement learning, integrating textual semantics with graph structure.
SAHM: benchmark and instruction-tuning dataset for Arabic financial NLP and Sharia-compliant reasoning with 14,380 examples.
Design principles for deploying machine learning models on extreme-edge AI hardware with latency/throughput constraints using spatial dataflow.
Audits fairness properties of LLM-based tabular classification on housing placement prediction with real nonprofit casenote data, evaluating multi-class classification error disparities.
Privacy-preserving entity alignment method for vertical federated learning without disclosing intersection information using noisy identifiers.
Evaluation framework for medical question answering systems using LLMs, assessing accuracy and health equity implications beyond semantic similarity.
Evaluation of LLM-generated obfuscated XSS payloads and their detection by machine learning-based security systems.
Method for detecting hallucinations in speech-based LLMs using attention map metrics without requiring gold-standard outputs.