Entdecken Sie sorgfältig ausgewählte Einblicke in die Welt der künstlichen Intelligenz
Greifen Sie auf vertrauenswürdige Nachrichten und aktuelle Branchenentwicklungen zu, die speziell für Entscheidungsträger aus Wirtschaft und Technologie zusammengestellt sind.

LLMs Interpret, Embeddings Organize, Graphs Emerge: Agent-Driven Compilation of Scientific Knowledge
von Shi-Ju Ran, Kun Zhang, Xi Wu, Liu-Si Yang, Wen-Jun Li am September 1, 2026 um 4:00 a.m.
arXiv:2608.29612v1 Announce Type: new Abstract: Sustained scientific work requires a knowledge substrate that carries interpretation across tasks and preserves paths to source evidence. We call this process \emph{scientific knowledge compilation} and implement it in ASKS, the \emph{Agent-Driven Scientific Knowledge System}. For each source, an LLM produces a readable Wiki view and machine-facing semantics. Deterministic checks convert the latter into a document-local GraphDelta, and embedding geometry together with explicit graph rules integrates the proposed […]
Spatial Matryoshka Training for Multi-Granularity Visual Document Retrieval
von Trishan Singha Roy, Arkadeep Acharya, Vishwajeet Kumar, Jaydeep Sen, Sachindra Joshi am September 1, 2026 um 4:00 a.m.
arXiv:2608.29951v1 Announce Type: new Abstract: Multi-modal late-interaction retrievers achieve strong retrieval on visually rich documents by representing each page as per patch embeddings and matching at the token level. However, this approach incurs high storage costs. Existing compression methods typically fix a single compression level at indexing time, limiting flexibility. We present ColSNAP (Spatial Nested Average Pooling)1, a training method that generates a nested hierarchy of compression levels directly from a backbone’s patch grid. By spatially […]
Action from Adjacent Set in Physical Space Outperforms the Best Prediction in World Models
von Liangyu Li, Qingwen Liu, Mingqing Liu, Wen Fang am September 1, 2026 um 4:00 a.m.
arXiv:2607.23602v2 Announce Type: replace-cross Abstract: Controllers based on sampling and latent world models assign a predicted terminal cost to each candidate action sequence, choose the minimum, execute its first action block, and replan. This rule can fail even when the terminal cost perfectly and accurately reflects the true task objective in the physical world. Residual prediction error can give an infeasible sequence an anomalously low cost, and a larger proposal pool gives such errors more chances to outrank feasible alternatives. We call this […]
FORESIGHT-9: Prospective and Process-Aware Evaluation of Adaptive Trading Agents
von Xiangxin Luo, Chengtian Hong, Haohua Li, Yongyi Xie am September 1, 2026 um 4:00 a.m.
arXiv:2608.29372v1 Announce Type: new Abstract: Retrospective backtests provide a limited test of adaptive trading agents: they cannot rule out historical contamination, expose sensitivity to a single realized market path, or reveal internal degeneration during long-horizon adaptation. We introduce FORESIGHT-9, a prospective and process-aware benchmark built from nine auditable counterfactual stress worldlines branching from a common July 2026 information boundary. Each worldline specifies staged macro-financial events and joint multi-asset anchors; a […]
GarmentWeaver: Schema-Aware Structured Synthesis for Multimodal Sewing Patterns
von Yinwen Lu, Weihao Luo, Yueqi Zhong am September 1, 2026 um 4:00 a.m.
arXiv:2608.30550v1 Announce Type: new Abstract: Multimodal Sewing pattern generation aims to infer executable sewing patterns from design cues such as sketches and textual descriptions. As an interpretable and simulation-compatible representation, sewing patterns are particularly valuable for digital garment creation. However, existing methods often model garment specifications as flat long sequences, which entangles garment structure with detailed parameters and leads to redundant components, inaccurate local details, and poor simulation compatibility. In this […]
CAER: Causal Action Effect Reweighting for World Model Training
von Jianjie Fang, Xvyuan Liu, Ziyou Wang, Rongze Tang, Zhaolu Wang, Zhuohang Li, Xin Zhang, Haisheng Su, Chen Gao, Wei Wu, Xinlei Chen, Yong Li am September 1, 2026 um 4:00 a.m.
arXiv:2608.30897v1 Announce Type: new Abstract: World models are becoming core infrastructure for embodied intelligence, with action-conditioned video generation providing controllable predictions of how scenes evolve after agent interventions. Yet existing models are commonly trained with space-time-uniform mean squared error, allowing abundant background tokens to dominate the gradient while sparse interaction dynamics remain under-optimized; such uniform fitting rewards reconstructing appearance rather than learning how actions change the world. We introduce […]
Collapsibility of Performance Metrics in Clinical Predictive AI
von Jo\~ao Matos, Ben Van Calster, Richard D. Riley, Paula Dhiman, Gary S. Collins am September 1, 2026 um 4:00 a.m.
arXiv:2608.30568v1 Announce Type: cross Abstract: Background: Population level assessments of predictive artificial intelligence (AI) can conceal performance disparities across subgroups. Fairness evaluations commonly rely on performance analyses across subgroups. However, some performance metrics are non-collapsible, meaning that the overall population performance value does not equal the weighted average of subgroup specific values. Objective: To examine the collapsibility properties of commonly reported performance metrics in predictive AI, with a focus on […]
Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents
von Hanlin Tian, Minhao Li, Yu Mi, Sihan Zhu, Zhao Yang, Yuxiang Wang, Hongquan Zhu, Qiufei Hu am September 1, 2026 um 4:00 a.m.
arXiv:2608.30322v1 Announce Type: new Abstract: Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks rarely control whether an agent has access to those conventions. We introduce a knowledge-gated task-construction protocol that separates a task instruction from a compact artefact containing private conventions, reference tables, and utility operators. Construction-time provenance, byte-identical task instructions across the provided- and withheld-artefact conditions, leak audits, and executable witnesses make […]
von Kiyan Rezaee am September 1, 2026 um 4:00 a.m.
arXiv:2608.28980v1 Announce Type: cross Abstract: Can the specialized architectures that machine learning has traditionally built for structured data be replaced by language-based models? This question is examined through a review of 159 papers (2016–2026) across nine modalities, with predictive accuracy considered alongside structural representation and computation. A distinction is made between performing a task and preserving and computing the structure that makes the task tractable, and existing approaches are organized into eight representational regimes, […]
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
von Anton Razzhigaev, Andrei Gritsaev, Andrei Kaznacheev, Nikita Dragunov, Roman Yampolskiy, Andrei Kuznetsov am September 1, 2026 um 4:00 a.m.
arXiv:2608.08311v3 Announce Type: replace-cross Abstract: We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evolution proceeds in two modes. In recursive free evolution, improvement is itself a task, and completing one evolution cycle can schedule the next. In experience-driven core evolution, ordinary work and social interaction expose bugs, rough edges, and inefficient context construction that lead to reviewed […]
Asymmetric Within-Document Predictive Learning for Scientific Document Representation
von You Zuo (ALMAnaCH), \’Eric de la Clergerie (ALMAnaCH), Beno\^it Sagot (ALMAnaCH) am September 1, 2026 um 4:00 a.m.
arXiv:2608.28625v1 Announce Type: cross Abstract: We study predictive pretraining for scientific document representation using the discourse structure of papers. We propose SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict method representations, and method representations are used to predict conclusion representations. Experiments on RELISH, high-influence citation, SciDocs, and cite prediction show that plain predictive training is viable but weaker than a […]
Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy
von Kaike Ping, Buse \c{C}ar{\i}k, Caleb Wohn, Xiaohan Ding, Tongshuai Wang, Eugenia Rho am September 1, 2026 um 4:00 a.m.
arXiv:2608.01017v2 Announce Type: replace-cross Abstract: Large language models can answer a medical question correctly and still abandon that answer when a user pushes back. We study this failure as medical sycophancy and ask when models are most likely to give in. Across five open-weight models, 500 MedQuAD questions, and 1.2 million trials, we use a fully crossed design over four conversational factors: user role, user evidence, interaction structure, and grounding. Medical sycophancy is nearly three times more common when users challenge an answer the model […]
Atom Learning Model (ALM): how a real classroom got tokenised
von Philipp Bogdan am September 1, 2026 um 4:00 a.m.
arXiv:2608.21106v3 Announce Type: replace-cross Abstract: The Atom Learning Model (ALM) tokenises a school curriculum. 757 pages of GCSE and Further Mathematics material were read by machine into 1,934 atoms, each one thing a learner can do in a single step, ordered by 4,616 machine-written prerequisite links. Both sides of a lesson are then expressed in that one structure: a question is a set of atoms plus everything beneath them, a child’s ability is a score between 0 and 1 on every atom of the same graph, and whether a question suits a child is arithmetic […]
Beyond Uncertainty: Multi-Solver Disagreement Rewards for Self-Evolving Reasoning Curricula
von Vinoth Selvendran, Zhanming Zhang am September 1, 2026 um 4:00 a.m.
arXiv:2608.30035v1 Announce Type: new Abstract: Self-evolving reasoning frameworks train a Challenger to generate questions exposing a Solver’s weaknesses, creating adaptive curricula without human data. However, existing approaches use a single solver’s sampling uncertainty as the Challenger’s reward. This creates a fundamental bottleneck: as the solver grows confident on the Challenger’s question distribution, all sampled answers converge identically, collapsing the reward to zero and starving the Challenger of learning signal. Critically, this single-model […]
Q-Strata: Hierarchical Bit Allocation for Mixed-Precision Quantization of Mixture-of-Experts LLMs
von Deokjae Lee, Sihun Chu, Hyun Oh Song am September 1, 2026 um 4:00 a.m.
arXiv:2608.30564v1 Announce Type: cross Abstract: Mixed-precision quantization (MPQ) assigns a different bitwidth to each linear layer of a large language model (LLM) to minimize the quantization-induced quality loss under a fixed budget, but Mixture-of-Experts (MoE) models contain these layers in every expert of every MoE block, so the allocation space grows far larger than in a dense model. Existing methods either allocate within each block under a uniform per-block budget, or allocate across blocks through an additive proxy, and neither directly optimizes a […]
FaVOR: LLM-Based Agentic Framework for Factor Mining via Empirical Validation
von Hyeonjin Kim, Minseok Kim, Seunghyeon Jung, Sujin Pyo, Huisu Jang, Woojin Lee am September 1, 2026 um 4:00 a.m.
arXiv:2608.30192v1 Announce Type: new Abstract: Traditional finance relies on experts to hand-craft factors through a principled process grounded in economic rationale. Recent LLM-based multi-agent systems have automated this process, scaling factor mining far beyond manual effort. However, these automated approaches optimize directly for returns and rarely check whether a generated factor still expresses the economic hypothesis that motivated it. We identify this inconsistency between mathematical form and economic meaning as a structural failure mode of […]
von Seth Grief-Albert, Jessica Bo, Difan Jiao, Ashton Anderson am September 1, 2026 um 4:00 a.m.
arXiv:2608.03044v2 Announce Type: replace-cross Abstract: Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak demographic sensitivity. We propose that much of this conflict stems from conflating two distinct tasks. We call the first task emulation, in which models generate individual responses that aggregate into a population distribution. We call the second task estimation, in which models […]
von Vipul Patel, Anirudh Deodhar, Dagnachew Birru am September 1, 2026 um 4:00 a.m.
arXiv:2608.30419v1 Announce Type: new Abstract: Healthcare workforce scheduling is an NP-hard optimization problem requiring simultaneous satisfaction of labor regulations, coverage requirements, employee preferences and cost objectives. Existing approaches (genetic algorithms, integer programming, constraint programming) model 6-12 constraints at shift-level granularity and cannot guarantee regulatory compliance. They also lack support for multi-role, multi-skill heterogeneity, mandatory break scheduling with midpoint control, acuity-weighted workload equity, […]
von Gaoming Zhang, Angqing Jiang, Jianchun Song, Kena Qi, Dayao Chen, Wei Lin, Defu Lian am September 1, 2026 um 4:00 a.m.
arXiv:2608.30553v1 Announce Type: cross Abstract: Generative Retrieval (GR) has emerged as a promising paradigm by mapping queries directly to Semantic IDs (SIDs) with powerful representation capabilities for candidate items. However, existing SIDs derived solely from item content create a semantic gap, failing to align dynamic query intents with static item representations. Furthermore, current generative paradigms rarely model user behavior sequences and are always bottlenecked by the high inference latency of beam-search autoregressive decoding. To address […]
ATLAS: Dual-Horizon Diagnostic Evaluation for Industrial Tool-Use Agents
von Wei Chen, Peilun Zhou, Zhaoyu Hu, Jiajun Chai, Zhongni Hou, Yufei Zhang, Derong Xu, Guojun Yin, Wei Lin, Zhi Zheng, Tong Xu am September 1, 2026 um 4:00 a.m.
arXiv:2608.30685v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed in user-facing services that require iterative tool use under dynamic business conditions. Reliable evaluation is essential for sustained improvement: it must reveal capability deficiencies, inform priorities, and assess interventions. Yet industrial agent service unfolds both through the iterative trajectory of a current request and through continued user interaction. Final-outcome assessment can therefore obscure where deficiencies arise and whether […]

