Stable Deep Reinforcement Learning via Isotropic Gaussian Representations
Deep reinforcement learning stability improvement through isotropic Gaussian embeddings under non-stationary training dynamics.
Deep reinforcement learning stability improvement through isotropic Gaussian embeddings under non-stationary training dynamics.
CeRA: improved parameter-efficient fine-tuning method that surpasses LoRA's linear constraints via manifold expansion with gating and dropout.
Analysis of transformer training trajectories under AdamW showing low-dimensional drift directions and batch-gradient alignment patterns.
Explainable AI method for highlighting token attributions in text classification using transformers.
Framework for autonomous neural architecture and hyperparameter search using self-evaluating RL agents without human supervision.
Research on robust policy training in partially observable reinforcement learning under adversarial latent state distribution shifts.
Systematic study of jailbreak attack scaling laws across LLM methods and model families using compute-bounded optimization framework.
Research on parameter-efficient fine-tuning for continual learning using representation-level optimization instead of weight-level black-box methods.
Zero-shot surgical duration prediction combining retrieval-augmented LLMs with Bayesian averaging for resource management.
Survey of privacy-preserving machine learning mechanisms for IoT devices covering federated learning and edge computing approaches.
Analysis of transformer training dynamics via Spectral Edge Dynamics, identifying coherent optimization directions vs stochastic noise.
Diffusion-based reinforcement learning policy using flow matching with direct entropy regularization and efficient gradient computation.
Detects hallucinations in virtually-stained histology using latent space analysis and neural precursor method.
Method for evaluating synthetic chest X-ray quality using embedded characteristic scores.
Universal sparse autoencoders for discovering and aligning interpretable concepts across multiple neural networks.
Multifidelity simulation-based inference framework for parameter estimation with expensive simulators.
Survey of AI-based methods for detecting and mitigating distributed denial-of-service attacks.
Ensemble of language models for automated tumor classification in cancer registry pathology reports.
Hardware-aware neural architecture search for encrypted traffic classification on IoT edge devices.
Testbed for evaluating AI reasoning with causal world models in low-data and out-of-distribution settings.
Universal distillation method for training efficient one-step generators from diffusion and flow models without GANs.
Improves one-step image generation from masked diffusion models using soft embeddings to enable gradient flow for fine-tuning.
Artificial Age Score framework modeling memory aging patterns in large language models across conversational contexts.
Lightweight disentangled concept bottleneck model for improved interpretability in neural networks.
Theoretical analysis of clipped gradient optimization under heavy-tailed noise with refined convergence bounds.
Novel synthetic data generation method for wireless network traffic forecasting to augment training datasets.
Multi-preconditioned LBFGS algorithm for training physics-informed neural networks using domain decomposition.
Study showing arbitrary token generation order in diffusion LLMs doesn't improve reasoning despite flexibility.
One-shot data augmentation technique for few-shot learning combining geometric perturbations with noise injection.
Federated learning approach over wireless channels using over-the-air aggregation without channel state information.
Theoretical and interpretability analysis of quantum extreme learning machines using Pauli-transfer matrix approach.
Method for efficient LLM reasoning balancing overthinking and underthinking in resource-constrained settings.
Learning-to-Defer framework extended to select experts and conditionally provide information like documents or tool outputs.
ML approach for predicting and discovering error patterns in vehicle diagnostic trouble codes using temporal sequence analysis.
Study of cultural biases in LLMs using author profiling from song lyrics in zero-shot settings across open-source models.
Curated resource list for OpenClaw AI system covering skills, plugins, MCP, tools, deployments and alternatives with editorial context.
Google AI Studio launches upgraded vibe coding with Firebase integration for building production apps via AI agents and prompts.
Opinion essay on AI exceeding human cognitive capabilities and societal implications; personal reflection without technical depth.
Case study of building AI-powered e-commerce platform from off-grid homestead using AI agents for content and automation.
HTTP API for Claude Code via MCP channels enabling programmatic agent control without terminal parsing; drop-in replacement for agentapi.
Anecdotal account of replacing Scrum team with AI agents for 10 days; content unclear due to mismatched description.
Open source local document parsing library designed for AI agents.
Distributed file store built on AWS CloudShell using Reed-Solomon erasure coding, AES-256-GCM encryption, and regional sharding for resilient storage.
Discussion of LLM/agent memory limitations and whether RAG/embeddings or local markdown files are more effective. Community exploration of context retention in agents.
GopherHole: open protocol for agent-to-agent communication across frameworks. Built-in memory, search, specialized agents. WebSocket-based, privacy-preserving.
FreeAgent: local-first AI agent with 60+ tools (image generation, OSINT, cybersecurity, automation). Runs entirely offline, no cloud dependency or subscriptions.
Analysis of when splitting AI prompts into subagents helps vs hurts output quality. Cognitive science perspective on prompt engineering and agent decomposition.
Intelligent Document Processing Leaderboard benchmarks 16+ LLMs/VLMs on document parsing, table extraction, OCR, and visual QA using 9,000+ real documents.
GlassWorm supply-chain malware campaign compromised 433 packages across GitHub, npm, and VSCode extensions; coordinated attack identified by security researchers.
Signal creator Moxie Marlinspike's Confer platform incorporating encryption into Meta AI systems. Privacy-focused AI infrastructure.