Privacy Attacks on Image AutoRegressive Models
Comprehensive privacy attack analysis on image autoregressive models, identifying membership inference and extraction vulnerabilities.
Comprehensive privacy attack analysis on image autoregressive models, identifying membership inference and extraction vulnerabilities.
Method for enforcing syntactic and semantic constraints in LLM decoding through MCTS-guided token-level control.
Large-scale corpus of 324,843 Python classes from open-source projects for training and evaluating LLMs on code generation.
RAG-based LLM workflow using domain-specific knowledge graph for automated single-cell type annotation in biology.
Study evaluating sparse autoencoders for detecting bugs in Java code, addressing software vulnerability detection.
Benchmark dataset (SealQA) for evaluating search-augmented LLMs on fact-seeking questions with conflicting or noisy search results.
Deployment case study of LLM-based platform for automated assessment of Romanian Bacalaureat exam questions using Gemini Flash.
Inference method for VideoLLMs that processes multiple frame subsets in parallel to improve temporal detail without increasing context window.
Technique to improve CLIP few-shot classification by addressing modality gap through semantic bridging between image and text embeddings.
Benchmark for evaluating LLMs on detecting demographic-targeted social biases across diverse content types and demographics.
Method to improve LLM performance in multi-turn conversations by reinforcing long-term planning and goal tracking through prompting.
Lightweight Disentangled Concept Bottleneck Model addressing bias in input-to-concept mapping for interpretable multimedia recognition.
Framework enabling diffusion models to adapt generation quality based on real-time network bandwidth constraints in cloud-to-device scenarios.
Minimax optimal algorithm for best arm identification under fixed sampling budget with applications to A/B testing.
Study of task transfer in Vision-Language Models examining how finetuning on one perception task affects performance on others.
Philosophical analysis arguing static value alignment approaches cannot ensure robust AI alignment under capability scaling and distribution shift.
OxEnsemble: Fair classification approach for low-data, imbalanced settings with demographic group constraints.
Study on selecting minimal training data subsets for example-based explanations of language model predictions using influence estimation.
Accordion-Thinking: Framework enabling LLMs to self-regulate reasoning step granularity through dynamic summarization for efficient inference.
Neuro-symbolic framework using differentiable logic programming to design and optimize quantum circuits.
Evaluation of 17 LLMs showing diagnostic reasoning degrades across multi-turn conversations compared to single-turn benchmarks.
HiCI: Hierarchical attention module for long-context language modeling, organizing information from local to global levels.
CodecSight optimizes streaming vision-language model inference by leveraging video codec signals for end-to-end efficiency.
DOVE benchmark for evaluating LLM cultural value alignment using open-ended generation, addressing limitations of multiple-choice formats.
Study investigating saturation points in recommender system performance as training dataset size increases, with reproducible Python implementation.
Research: Chain-of-thought models generate 52-88% of tokens after answers are already recoverable, revealing inefficiency in reasoning.
Brief mention of LLM routers injecting malicious tool calls as security issue. Insufficient detail.
Analysis of AI agents as SaaS replacement with integrated database, logic, and UI.
Linter tool for AI agent context files and MCP configs. Detects stale references, token waste, hardcoded secrets. Cites research on context bloat reducing agent performance.
Open-source platform for managing AI coding agents as teammates. Agents autonomously handle task assignment, code writing, and progress tracking without prompt copying.
Research on video generation models struggling with multi-subject action binding. Introduces ActionParty with per-subject state tokens to improve action accuracy.
Study reports Google's Gemini models generate inaccurate search results 9-15% of the time across 8,652 queries. Evaluates LLM reliability in production systems.
Facebook Marketplace MCP integration allowing Claude to search and monitor deals via command line. Demonstrates AI agent tool use.
Discussion question about managing token costs and data exposure in production AI agent systems. Operational concern without detailed analysis.
arXiv framework page for collaborative feature development. No technical content about the claimed compiler-LLM optimization research.
Tend: lightweight infrastructure for managing multiple concurrent AI agents and projects. Handles agent coordination, context switching, and multi-project workflows.
Five LLM-based agents play Texas Hold'em with distinct personalities and reasoning styles.
Discussion of LLM application to code patch review. Limited content, but relevant to developer tools and LLM applications.
Tool for streaming terminal activity to observe AI agent execution. Visualization/monitoring for agent workflows.
Open-source local sandbox environment for AI agents on macOS/Linux. SDK/CLI enables persistent state, file I/O, pause/resume, and desktop environment interaction for agent workflows.
AI resources for clinical workflows: evidence search, guideline reconciliation, and documentation support in healthcare.
Resources for evaluating, deploying, and scaling AI in regulated financial services environments.
Overview of OpenAI's evolution from research to consumer and developer products for AI applications.
Trinity-Large-Thinking: open-source frontier reasoning model for long-horizon agents and tool calling. Apache 2.0 licensed weights on Hugging Face. Supports complex multi-turn interactions.
AnimTOON: Model for generating Lottie animations from text with 3-4x fewer tokens than alternatives, supporting skeletal animation.
Research on using AI to synthesize workload-specific OLAP database engines optimized for specific query patterns.
Lingle is a voice-based AI agent that simulates personal language tutoring via conversation, addressing flexibility and pricing issues with traditional platforms.
Memoriki: Template combining LLM Wiki pattern with MemPalace MCP server for persistent personal knowledge bases with semantic search and temporal graphs.
Agent Tuning technique using recursion to achieve predictable Claude Code agent output through self-observation and instruction editing.
Analysis of open source supply chain attacks in March 2026 including credential theft and malicious library poisoning.