AI/ML papers collected from arXiv, OpenAlex, Semantic Scholar, and bioRxiv. Recommended ranking blends AI importance (60%) with citation and recency popularity (40%).
AIThis paper introduces 'The AI Scientist', a pipeline that automates the entire scientific research lifecycle. The system generates ideas, writes code, runs experiments, analyzes data, writes the full manuscript, and performs its own peer review, with the generated manuscript passing the first round of peer review for a top-tier machine learning conference workshop.
AIThis paper systematically organizes the latest techniques in LLM across four dimensions: pre-training, post-training, utilization, and evaluation. It also identifies key research issues and challenges such as theoretical foundations, efficient scaling, alignment, and agentic capabilities.
AIAxiom is an RNN that constructs a unitary transition matrix via a product of Householder reflections, enabling lossless processing of sequences of arbitrary length. It achieves 76.5-99.9% accuracy on the delayed copy task with 13x fewer parameters than LSTM, and when attached to GPT-2, it retrieves facts from 7 chunks back with 62.3% accuracy.
AIWe prove in four stages that the token embedding layer is the geometric foundation of transformer attention. Fourier-based training techniques (PFFT, FGP) and No-Q attention preserve the embedding's geometric authority and achieve lightweight performance.
AIWe propose Cognitive Impedance Matching Theory (CIMT), a compiler theory that enhances the capabilities of a fixed LLM model through world-side interface, validation, repair, and audit design. Human evaluators or LLM judges are not treated as privileged evaluators but modeled as fallible measurement channels, providing a framework to prove system reliability using only observable atomic data.
AICo-Scientist is a multi-agent AI system built on Gemini that assists scientists by generating and refining novel research hypotheses based on their objectives and existing evidence. It uses an asynchronous task execution framework and a tournament evolution process to improve hypothesis quality, demonstrating practical value in biomedical applications like drug repurposing.
AIA field experiment with 758 knowledge workers at a global consulting firm examined GPT-4's impact on task performance. The study found that while AI significantly boosted productivity and quality on most tasks, it hindered performance on a complex managerial task, revealing a 'jagged technological frontier'.
AIThis paper presents an interface theory that distinguishes genuine improvement from illusion in AI systems that change their own evaluation criteria. It develops a mathematical framework ensuring the stability of self-improvement loops through replayable observable interfaces and certified stable gains.
AIThis study quantitatively evaluates two specialized clinical AI tools against three frontier general-purpose LLMs across three medical benchmarks. The frontier LLMs outperformed the clinical AI tools in all evaluations, which performed comparably to an automated Google Search AI Overview on the real clinical queries benchmark.
AIThis study analyzes the practical effectiveness of differential privacy (DP) guarantees when adapting large language models (LLM) for sensitive applications. The researchers benchmark privacy risks using state-of-the-art attacks while systematically varying the adaptation data distribution, identifying key factors for achieving practical privacy protection.
AIProposes the Spiral Time framework, which represents time as 2D spiral coordinates (radius=trend, angle=seasonality). In LSTM and Transformer experiments, it improves MAPE by 83% over scalar time, with monotonic performance gains across all experiments.
AIThis paper provides a comprehensive survey of Multimodal Large Language Models (MLLMs) focused on vision-language tasks such as image captioning and visual question answering. It examines MLLM architectures, training pipelines, and practical applications, while highlighting fundamental constraints like information bottlenecks and data-processing limits.
AIQwen-Scope open-sources 14 groups of sparse autoencoders (SAEs) for Qwen3/3.5 series models and demonstrates their use in inference-time steering, evaluation analysis, data-centric workflows, and post-training optimization, showing that SAEs can serve as practical interfaces for model development beyond post-hoc analysis.
AIThis study analyzed 7.3 million academic articles from 2020 to 2025 to track the widespread adoption of large language models (LLMs) in scholarly research. The analysis found that by 2025, an estimated 57% of published articles showed evidence of LLM influence, with significant variation in adoption rates across regions, institutional ranks, and academic disciplines.
AIThis research identifies necessary fixes to scale Schedule-Free Learning for large language models. The proposed ScheduleFree+ method significantly outperforms Warmup-Stable-Decay schedules, especially in long-duration training.
AIThis research proposes 'natural identifiers (NIDs)' to address the challenges of privacy auditing for large language models (LLMs). NIDs are structured random strings naturally occurring in training datasets, enabling post-hoc differential privacy auditing without retraining and dataset inference without needing a private held-out dataset.
AIThis study proposes a variational recurrent neural network (VRNN) that integrates latent random variables into the hidden state of a recurrent neural network (RNN). Experiments on speech and handwriting data demonstrate that VRNN better models the variability of structured sequential data than existing models.
AIWe consider a partial observability model in bandit problems where the learner can observe losses of some other actions in addition to its own. We propose the first algorithm that guarantees near-optimal regret without prior knowledge of the observation system, and show that the implicit exploration strategy is more efficient computationally and information-theoretically than existing methods.
AITo address the difficulty of model selection in deep unsupervised domain adaptation (Deep UDA), we propose Deep Embedded Validation (DEV), which embeds adapted feature representations into the validation procedure to provide an unbiased estimator of the target domain risk. Variance is reduced using the control variate technique, and effectiveness is demonstrated theoretically and empirically.
AIChatlaw is a multi-agent legal assistant specialized in the Chinese legal system, using a Role-Aligned Mixture-of-Experts (RA-MoE) architecture. It achieves a 7.73% improvement in accuracy over GPT-4 on the LawBench benchmark and an 11-point higher score on the legal professional exam.
AIThis paper systematically categorizes and analyzes graph neural network (GNN) models that address heterophily in graphs. It summarizes various approaches to overcome the limitations of existing GNNs that assume homophily and suggests future research directions for heterophilic graph learning.
AIThis paper presents MicroGrowAgents, an AI-driven, agent-based system that automates the design of optimized microbial growth media by integrating knowledge graphs, metabolic modeling, and optimal experimental design. The system uses specialized agents to query biological knowledge, mine literature, and generate statistically optimal experiments, aiming to reduce experimental burden and accelerate the discovery of growth-promoting conditions.
AIThis paper proposes that the universe operates as a necessity system, where organizational constraint drives structural configurations toward the golden ratio through a recursive mechanism identical to the Fibonacci sequence. The theory aims to explain the cosmological constant without fine-tuning and presents five independent empirical pillars of support, including nuclear morphometry and high-redshift galaxy observations.
AIReveals that Multi-modal Large Language Models (MLLMs) risk leaking sensitive information embedded in images, and presents a comprehensive dataset MM-Privacy for evaluating these risks. Experiments confirm that various MLLMs expose personal information across different tasks, and task inconsistency increases privacy risks.
AIThis study examines the impact of weakened patent protection following the Alice Corp. vs. CLS Bank decision on firms' innovation, competition, acquisitions, lawsuits, and employment agreements. It uses large language models (LLMs) to identify the potential exposure of firms' patent portfolios to the Alice decision and investigates the resulting unequal impacts.
AIAccurate carbon emission forecasting in power distribution networks is a critical challenge due to the integration of large-scale electric vehicles and renewable energy sources. This paper proposes CarbonGPT, a model that uses a causal encoder and a meta causal graph dictionary to address spurious correlations and enhance LLM-based prediction.
AIThis survey addresses the multilayered challenges of deploying large language models on edge hardware, including compression, compiler behavior, and system-level trade-offs. It provides a deployment-centric taxonomy of compression strategies, analyzing their interaction with hardware toolchains and highlighting limitations in current benchmarking suites.
AIThis survey systematically analyzes methods for integrating external knowledge into LLMs to address hallucinations and knowledge gaps. It discusses parametric and non-parametric approaches to improve reasoning and factual accuracy in domain-specific tasks.
AIThis research analyzes which brands are recommended within specific categories and the concentration of their ownership across large language models. The authors propose three exploratory metrics—Category Ownership Index (COI), Competitive Vacuum Index (CVI), and Displacement Score (DS)—and conduct an empirical analysis across 3 models, 5 industries, and 250 category queries, finding results that challenge a strong winner-takes-all narrative.
AIThis paper proposes BRIDGE, a benchmark for evaluating large language models on understanding real-world clinical practice texts. The benchmark includes various clinical tasks such as disease diagnosis, treatment planning, and patient status summarization.
AIThis paper presents a large-scale empirical study evaluating ten LLMs with seven prompting strategies against nine traditional techniques for software vulnerability analysis. The study finds that existing prompting strategies often lead to LLMs underperforming traditional methods, and proposes a vulnerability-specific chain-of-thought prompting (VSP) to improve performance.
AIThis paper provides a systematic survey of how Large Language Models (LLMs) are utilized to address Operations Research (OR) problems. It analyzes the roles of LLMs in OR, such as model formulation, algorithm design, and solution verification, along with practical applications and benchmark datasets.
AIThe human brain represents sentence meaning differently depending on word order, while large language models (LLMs) are less sensitive to order and rely more on context. fMRI experiments and model analysis revealed differences in sentence processing between humans and LLMs.
AIThis study analyzes how state-level media control influences the outputs of large language models. It reveals correlations between internet censorship and media control indicators across countries and biases in LLM outputs.
AIThis is the first survey paper to comprehensively examine the pretraining data exposure problem in LLMs from the perspectives of data contamination and membership inference. It systematizes attack and defense methodologies and suggests future research directions for evaluation integrity and privacy protection.
AIThis paper analyzes how the capabilities of Large Language Models (LLMs) emerge through the lens of emergence concepts from complexity science. The study reviews several approaches to quantifying emergence and questions whether LLMs possess emergent intelligence.
AIOCR and multilingual text understanding are major failure modes of multimodal LLMs. The proposed framework combines synthetic data generation, OCR-aware fine-tuning with LoRA, and visual chain-of-thought prompting to significantly improve OCR completeness and multilingual translation accuracy in complex real-world images.
AIA comprehensive survey analyzing architectures, evaluation methods, and safety of LLM-based network operations and AIOps agents. It organizes related research around autonomy hierarchy, tool scope, evidence traces, and assurance contracts, emphasizing the need for workflow-centered evaluation beyond static QA.
AIWe propose a method for LLM-based probability estimation that introduces hierarchical factor structures and causal Bayesian networks to reduce 'unknown' predictions in sparse factor spaces. Experiments show that compared to direct LLM inference, it significantly reduces unknown predictions, provides more reliable probabilities, and reduces time and token costs.
AIThis paper traces how an organization's declared purpose is translated into the criteria, structures, and signals that govern real decisions. It demonstrates through case studies that meaning is systematically reformulated as it moves through governance systems.
Version 2 (2026-08-04). Revised to the NLL Universal Paper Format v6. The original text is retained in full; nothing has been deleted. Corrections appear as marked blocks placed at the section that carries the claim, and each one states what the paper said, what the data show, th…
AIThis paper systematically reviews empirical studies on the application of Large Language Models (LLMs) in education. The studies primarily apply LLMs to student support, teacher assistance, and automated assessment, reporting benefits such as improved learning outcomes and personalized feedback. However, the review also highlights significant challenges, including hallucination, bias, potential for academic dishonesty, and privacy concerns.
AIThe Encyclotron is a reproducible instrument that quantitatively measures the degree to which AI summarization systems distort or simplify scholarly knowledge. It calculates the gap between scholarly and retrieval graphs using variables such as compression loss, invention, and distortion, and tracks changes over time.
AIThis study proposes a novel strategy combining transcriptomic data and machine learning to predict the function of oxidative phosphorylation (OXPHOS) genes in C. elegans. By integrating supervised learning ensembles and cluster-based inference, it identifies promising new candidate OXPHOS genes that were previously unannotated.
AINiCLIP is a contrastive language-image pretrained model trained on over 23,000 neuroscientific articles to predict cognitive tasks, concepts, and domains from brain activation patterns. Performance is optimized with full-text articles and a curated cognitive ontology, showing accurate predictions on group-level activation maps but limitations on noisy subject-level maps.
AIThis study developed a machine learning framework to automatically recognize and quantify multiple features of axons and myelin from electron microscopy images. When applied to spinal cord fibers in variably hypomyelinated mice, it demonstrated that reductions in myelin sheath thickness and length correlate with changes in mitochondrial density and periaxonal area.
AIWe introduce the concept of weighted rules under the stable model semantics following the log-linear models of Markov Logic. This enables resolving inconsistencies, ranking stable models, assigning probabilities, and applying statistical inference.
AIProposes a new theoretical framework called Semantic Physics, exploring the convergence horizon of information theory and semantics through concepts of semantic saturation and ontology competition. Introduces original concepts such as compression survival and semantic dark matter to analyze the limits of semantic processing in AI systems.
AIThis paper proposes a new approach using graph data to address the lack of spatial reasoning abilities in LLMs. It envisions a future where search engines integrate with LLMs to answer complex spatial questions through graph-enhanced reasoning for domains like urban planning and civil engineering.
AIRapid decision-making reduces the time to detect discrepancies between intent and actual behavior, allowing errors to accumulate. This paper introduces the concept of 'Translation Half-Life' to measure how quickly interpretive shifts become embedded in an organization before they can be corrected.
AIPresents an ontology defining information as the structural pattern of energy differences (energy texture). Through the six pillars of Energy-Efficiency Theory (EET), it explains why information is inherently constrained by energy.
AIThe CHOPSTICK architecture is a framework for the AGI era that treats AI outputs only as candidate materials subject to human review and approval, separating human discretion and responsibility. It redefines human intelligence not as a subordinate form of AGI but as a human-centered discretionary structure, preventing AI outputs from being mistaken for human judgment or decisions.
AIProposes the Provenance Erasure Rate (PER) metric to measure the proportion of claims in AI-composed outputs that lack explicit attribution to original sources. PER focuses on source visibility rather than truthfulness, and a case study of a Google AI Overview that fabricated a false biography from real poetry fragments showed PER=1.0.
AIDeepMind's 'AI Agent Traps' paper classifies adversarial influence on agents into six categories, but this is merely 'meaning feudalism' that presupposes platform sovereignty and treats external influence as attack. The paper omits legitimate environmental influence (commons repair), and the author proposes S4 (Legitimate Influence Blindness) as a new shadow.
AIMeaning changes leave traceable traces before they appear in performance indicators. This paper derives Translation Coherence as a measurable property from governance artefacts and allocation patterns, making alignment empirically observable.
AITo address the problem of meaning degradation when AI accelerates decision-making, this paper proposes a closed-loop architecture that preserves human interpretive control while supporting AI analysis. By constraining how intent is translated into criteria, metrics, and allocation rules, it prevents drift within the system and enables traceable decisions.
Abstract Accurate medical image segmentation is critical for early medical diagnosis. Most existing methods are based on U-shape structure and use element-wise addition or concatenation to fuse different level features progressively in decoder. However, both the two operations ea…
AIProposes the concept of the 'Operating Spine' as a minimal causal architecture that can identify misalignment between intent, decision criteria, and outcomes within an organization's internal structure. Instead of post-hoc analysis, it makes drift observable within the decision system.
AIIn the AGI era, a structural methodology for multi-layer cross-verification of AI outputs before linking them to roles, responsibilities, evidence, etc. The number and composition of verification layers vary according to output type, domain risk, role sensitivity, etc.
This paper reports on a review of a number of well-known naturalistic decision making. NDM, models (Zsambok & Klein, 1997). Key features of these models were identified and were found to represent different views of the same naturalistic decision making process. The key features…
AIThis study presents the first systematic benchmark of four structural MRI foundation models on sex classification, brain age prediction, and Parkinson's disease classification. 3D-Neuro-SimCLR showed the most consistent performance overall, but all models failed to classify early-stage Parkinson's disease above chance.
AIThis research clarifies the evaluation factors for image steganographic algorithms, focusing on the widely used LSB technique. It analyzes the effectiveness of steganography based on three main parameters: payload capacity, image quality measure, and security measure.
ABSTRACT The concept of the intelligent agent represents a long‐standing pursuit in artificial intelligence. Recent breakthroughs in large language models (LLMs) have catalyzed a paradigm shift, enabling the development of sophisticated agents that exhibit advanced reasoning, pla…
AIThis working paper introduces the BIFACE-Based Sentence Coordinate Documents (SCD) framework, where human-readable surfaces and AI/AGI-referable coordinate layers coexist. SCD is a pure coordinate-reference grammar, not a storage, execution, approval, or legal-effect system, providing a foundation for subsequent papers.
IQ-TREE (https://iqtree.github.io/) is a widely used open-source software tool for efficiently inferring phylogenetic trees under maximum likelihood. Here, we present IQ-TREE version 3, the third major release of the software. IQ-TREE 3 significantly extends version 2 with new fe…
AIThis paper proposes a structural governance framework that positions AI+AGI outputs as reference and assistive outputs rather than final decisions or evidence. It takes a non-executable approach to limit legal and liability effects through coordinate referability and boundary setting under human discretion.
AILaBB-CAT is an open-source tool for storing and automatically annotating speech corpus transcripts. It uses an annotation graph data model to support flexible annotation at various levels of granularity.
AIThis paper proposes an HTS (History Time-based Sealing) framework for time-based sealing of AI output records, human discretion records, role-based responsibility references, etc., in the AGI era. HTS does not automatically confirm evidence or create legal effects; it merely preserves the temporal reference of output records as a time reference layer.
AIProposes a technical specification defining seven components (JSON-LD definition, disambiguation matrix, keyword block, negative tags, semantic integrity markers, DOI reference list, evidence membrane) for entity-level retrieval. Differentiates from existing metadata standards, with a worked example using the Lee Sharks knowledge graph.
AIAll semantic operations are compression operations, and the key variable is what the compression burns and where the unrecoverable cost lands. Three regimes are defined: lossy, predatory, and witness compression, and a transfer law connecting semantic economy and semantic physics is proposed.
The previous papers in this series established two propositions. The first is that verification — the capacity to determine whether intelligence corresponds to reality — is becoming the scarce resource of the intelligence age. The second is that intelligence systems operating wit…
AIThis paper interprets light as the extreme manifestation of free-state energy from the perspective of Energy-Efficiency Theory (EET). It takes the speed of light as a measurement benchmark, explains light-matter interaction as an energy-efficiency cycle at the quantum level, and views vision as an extension into cognition.
Large Language Models (LLMs) have achieved prominent success in various applications, driven by their foundational capabilities and strong generalization potential. While their impact on natural language processing is well-established, recent works highlight their significant pro…
AIThis paper establishes the constitutional ontology of gravitation within the Energy-Efficiency Theory (EET) framework. Gravitation is defined not as a fundamental force but as the long-range gradient of the total constrained potential across the constraint network, and the gravitational constant is upgraded to a structural identity derivable from constraint network parameters.
Machine learning (ML) methods for network anomaly detection are emerging as effective proactive strategies in threat hunting, substantially reducing the time required for threat detection and response. However, the challenges in training and maintaining ML models, coupled with fr…
AIThis study proposes a method to teach metric relations in right triangles using a grid. It presents a new teaching approach that deduces formulas without relying on triangle similarity.
Ontology engineering (OE) is a complex task in knowledge representation that relies heavily on domain experts to accurately define concepts and precise relationships in a domain of interest, as well as to maintain logical consistency throughout the resultant ontology. Recent adva…
This working paper introduces Consent and Order Candidate Layers as a non-executable framework for AI+AGI-generated consent phrases, order phrases, approval phrases, payment request phrases, contract phrases, and action request phrases in the AGI era. Building on Papers 1 through…
This paper develops provider-independent structural reference layers for AI and AGI environments. It addresses the risk that documents, roles, authority conditions, state histories, seal references, human discretion, economic references, preservation contexts, and propagation bou…
Introduction: Long-term time series forecasting (LTSF) has gained significant attention in recent years. While various specialized designs exist for capturing temporal dependency, recent studies have shown that even a single linear layer can achieve competitive performance. This…
As a technique that can compactly represent complex patterns, machine learning has significant potential for predictive inference in multi-source and heterogeneous information fusion scenarios. K-fold cross-validation (CV) is the most common approach for ascertaining the likeliho…
Abstract Self-lubricating bearings are widely used in aerospace, marine, and other fields due to their excellent performance. Accurate wear prediction for self-lubricating bearings is crucial for ensuring reliability and safety. However, achieving both physical interpretability a…
Large language models are increasingly deployed across mobile and edge environments, where privacy-sensitive and heterogeneous user data raise critical concerns of copyright infringement, data leakage, and regulatory non-compliance. Machine unlearning has thus emerged as an essen…
3D Gaussian splatting (GS) has emerged as a transformative technique in radiance fields. Unlike mainstream implicit neural models, 3D GS uses millions of learnable 3D Gaussians for an explicit scene representation. Paired with a differentiable rendering algorithm, this approach a…
What is life, physically? This document provides the constitutional answer: life is the active maintenance of a constraint network --- a configuration of matter that continuously invests free-state energy ($\dot{E}_{\mathrm{main}} > 0$) to preserve its own structure against spont…
We propose a deterministic structural semiotics framework that treats residual evolution as an interpretable signal rather than noise, enabling inference to proceed directly from structured deviation without requiring exhaustive prior specification of system behaviors. This paper…
Deep learning models often encounter two key challenges in developing intelligent and scalable forecasting frameworks for renewable energy systems: input feature space dimensionality and sensitivity to hyperparameter settings. These limitations increase computational cost and com…
The DSFB Structural Semiotics Engine [1] is a domain-agnostic deterministic framework forresidual-based structural interpretation in dynamic systems. It treats residual trajectories asprimary inferential objects — reading drift direction, slew magnitude, and admissibility gram-ma…
Role classification involves grouping hosts into related roles. It exposes the logical structure of a network, simplifies network management tasks such as policy checking and network segmentation, and can be used to improve the accuracy of network monitoring and analysis algorith…
Large language models offer new opportunities for behavioural science, but their rapid evolution poses challenges for research rigour. We introduce a consensus-based reporting checklist to improve transparency, reproducibility and ethical accountability of large-language-model-ba…
The public release of large language models (LLMs) in late 2022 has fundamentally altered the landscape of scholarly medical publishing. LLMs now permeate every stage of the academic publishing pipeline, from manuscript drafting and peer review to editorial decision-making, with…
Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis, existing benchmarks largely reduce OCT understanding…
To address the common problem of noisy label interference in text classification tasks, this paper proposes a selfcorrecting text classification method for large language models to overcome the shortcomings of semantic learning shift, unstable category discrimination, and accumul…
The application of Large Language Models (LLMs) in Model-Driven Engineering (MDE) has emerged as a rapidly evolving research area. While existing systematic literature reviews have examined specific technical approaches, a comprehensive mapping of the broader research landscape (…
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualizat…
Integrating external tools with Large Language Models (LLMs) has emerged as a promising paradigm for accomplishing complex tasks. Since LLMs still struggle to effectively manage large tool collections, researchers have begun exploring retrieval-based methods to pre-select the mos…
In multimodal large language models (MLLMs), inference cost is largely dominated by the visual token prefix rather than the language backbone, making token reduction a key factor for improving efficiency. Existing approaches typically assign independent importance scores to visua…
Large language models are increasingly deployed on device and at the edge, where memory capacity, bandwidth, latency, and privacy requirements dominate system behavior. This survey systematizes the end side stack from algorithms to systems. On the model side, we present a clear t…
Current large language models (LLMs) can work with structured information and even assist developing program code, but can they support working with knowledge graphs (KGs) as well? Which LLM is offering the best capabilities in the field of semantic web and knowledge graph engine…
Large language models (LLMs) have demonstrated remarkable reasoning and generation capabilities in various natural language tasks. However, they often struggle with hallucinations or reasoning errors, particularly when handling domain-specific knowledge or complex multi-hop reaso…
Significance Understanding why people make the choices they do is central to decision science. We show that large language models can uncover people’s stated reasons from free-text reports, achieving 95% alignment between actual choices and those implied by the identified reasons…
The transition to large language models (LLMs) represents a major shift from traditional machine learning (ML) modeling. it raises new requirements for the technological systems supporting model development and deployment. Organizations that have built robust healthcare systems w…
Large language models increasingly need to accumulate and reuse historical information in long-term assistants and agent systems. Simply expanding the context window is costly and often fails to ensure effective context utilization. We propose $\delta$-mem, a lightweight memory m…
Conventional signature-based defenses no longer protect the heterogeneous, large-scale infrastructures that the Internet of Things (IoT) now constitutes. Large language models (LLMs) and agentic artificial intelligence (AI)—systems that autonomously perceive, reason, plan, and ac…
Psychological instruments designed for humans are increasingly used to assign large language models (LLMs) stable psychological profiles that affect their usability, safety assessment, and use as proxies for human participants in research. Using a formal psychometric framework, w…
Why does full-pipeline FP4 training of large language models often diverge, even when forward activations and activation gradients remain stable? We address this question through a controlled study of MXFP4 quantization in transformer training, progressively enabling FP4 across f…
Jailbreak prompts can bypass alignment guardrails in large language models (LLMs) and elicit unsafe outputs, making reliable deployment-time detection critical. Prior detection approaches largely rely on a fixed metric space, e.g., raw inputs, gradients, or hidden features, in wh…
Large language models (LLMs) have emerged as powerful foundation models with strong reasoning capabilities across domains. Beyond reactive text generation, agentic LLMs enable autonomous workflow execution through modular task decomposition and coordinated tool use. In structural…
Large language models remain vulnerable to jailbreak backdoor attacks, where adversaries poison safety alignment data to embed hidden triggers that bypass safety mechanisms. Existing defenses often require comprehensive attack information or multiple triggered examples, making th…
multiobjective evolutionary algorithms (MOEAs) have achieved notable success in recommendation systems (RS) by meeting diverse user needs. However, existing MOEAs lack effective methods to coordinate the challenges of cold start, low convergence of multiple objectives and lack of…
Editor’s notes: This article presents a unified, resource-efficient framework that leverages large language models to jointly automate RTL generation and SystemVerilog assertion synthesis, overcoming long-standing fragmentation between design and verification. By fine-tuning LLaM…
Abstract Large language models (LLMs) are increasingly deployed to support human decision-making. This use of LLMs has concerning implications, especially when their prescriptions affect the welfare of others. To gauge how LLMs make social decisions, we explore whether five leadi…
Smart manufacturing relies on programmable logic controllers (PLCs) that translate sensor inputs into actuator commands. Generating PLC programs in legacy textual languages such as Mitsubishi FX-series Instruction List (IL) remains an expert-only task, and IL’s deprecation in IEC…
Penetration testing is essential to securing modern web infrastructures, yet traditional manual methods struggle to keep pace with their scale and complexity. Large Language Models (LLMs) offer new opportunities for automating these tasks, but existing approaches face two persist…