October 7th

Session 1 (9:30-10:15AM)

 

Design and analysis of experiments assessing handheld RFID readers in complex environments

Justin Strait, LANL

Abstract: Commercial Radio Frequency Identification (RFID) inventory tracking systems have been successfully deployed in numerous industries, but their viability for complex nuclear material environments is less understood. Success of RFID in this space requires thorough examination of factors which influence its ability to inventory random collections of objects, without much prior knowledge about the objects’ properties. In this work, we design and analyze experiments to identify relevant handheld RFID settings which optimize performance for one such challenging scenario, a typical cluttered shelf configuration comprised of numerous metal containers. We specify a full factorial split-plot design, and model the probability of a successful match for each container through a hierarchical generalized linear mixed model. The model is used to infer factors which maximize match probabilities for individual containers and the full configuration, and examine sensitivity of performance to these factors. Crucially, we model match probabilities without explicitly accounting for specific covariates about the containers themselves, in order to mimic practical scenarios in which information about objects being inventoried by RFID is unavailable. 

 

The perfect is the enemy of the good

Bradley Jones

Abstract: Historically, orthogonal arrays (OA) have been considered the perfect experiment design. This is because for a given number of runs an OA minimizes the variance of prediction of all the parameters to be estimated. Also, OAs are easy to analyze requiring only the computation of means and variances to obtain model’s parameter estimates and standard errors.


This talk considers the many ways that for DOE the perfect design (i.e. an OA) can fail to meet the requirements of a specific experimental setting. In the past in such cases, experimenters have sometimes chosen to change the requirements of the problem in order to use an OA (or other textbook design). This talk advocates, rather, for using a design that matches the requirements of the problem despite losing a modest amount of efficiency in estimating the parameters of the model. The key is to fit the design to the problem. Do not change the problem to fit the design.

 

Coming Soon

Shihao Yang, Georgia Institute of Technology

Abstract: Coming Soon

Session 2 (10:30AM-12PM)

 

Use of Split-Plot Optimal Design for Optimization of a Constrained Two-Step Chemical Reaction Process: A Case Study Example

Richard Montes, Alnylam Pharmaceuticals, Inc.

Abstract: Process development is a requirement in chemical pharmaceutical manufacturing. Design of experiments (DOE) is an essential tool for an empirical approach to process development. Most real-world processes have design constraints that prevent the use of classical experimental designs. Algorithmically-generated optimal design (OD) is particularly useful to accommodate such constraints with a minimal number of experimental runs. DOE application to the optimization of a two-step chemical reaction process with logistical and chemistry constraints is exemplified in this paper. A 5-factor split-plot OD is used, requiring mixed-effects model fitting of the product and impurity responses. The complexities in the analysis introduced in accommodating the constraints are considered. Statistical models generated from the DOE predicted results of confirmatory runs, demonstrating the adequacy of the models. The DOE study successfully identified the optimal process condition in a systematic and efficient manner. The synergy achieved when chemists and statisticians partner together is key to this success. Case studies like this paper aim to educate practitioners on the use of OD to promote its more widespread use.

 

Powerful Experimental Designs for Multiple Hypothesis Tests

Jonathan Stallrich, North Carolina State University

Abstract: Classical optimal design criteria such as A-, D-, and E-optimality target minimizing variances in an overall sense. Surprisingly, this objective may not directly target maximizing power for hypothesis tests with multiple degrees of freedom. Power is governed by the noncentrality parameter, which, unlike variances, is (1) invariant to a wide class of transformations of the null hypothesis and (2) depends on the unknown mean model parameters. This talk develops a weighted optimality framework that targets maximizing power for grouped hypotheses which may have differential importance. We propose weighted E-, A-, and D-optimality criteria with respect to a hypothesis-based information matrix which we show maximize the noncentrality parameters in an overalls sense. Exchange algorithms are used to construct exact designs in discrete spaces, with examples illustrating how criterion choice and weighting affect design structure in both standard and constrained settings. One such example involves an aerospace application in which hypothesis testing criteria are employed to address a challenging model constraint that forces linear dependences in the model matrix.

 

Sample size and design optimization in the presence of transition costs

Peter Goos, KU Leuven

Abstract: Many industrial experiments involve hard-to-change factors which require substantial time to change. In such scenarios, completely randomized designs are only feasible if the number of runs is small. By foregoing complete randomization, like in split-plot, split-split-plot, and staggered-level designs, the number of runs may be increased. However, the literature provides no guidance concerning which of these restricted-randomization designs is best for any given experiment. It is even not sure that the best design for a given problem is one of these well-studied designs.  Our goal is to present an algorithm to construct optimal restricted-randomization designs directly taking into account the budget constraint and difficulty of resetting the individual factors. Each factor is assigned a transition cost, where more difficult to change factors are assigned a higher cost. The algorithm contains two innovative aspects. First, the total number of runs is not fixed in advance. It is part of the optimization procedure. Second, the algorithm is not constrained to well-known restricted-randomization structures, but rather considers a broader variety and will therefore generally produce a better design. An additional benefit is that the planner of the experiment no longer requires knowledge on randomization structures, thus making optimal experimental design more accessible in the presence of hard-to-change factors.

 

Data Quality: Missed Opportunity or Golden Opportunity?

Roger Hoerl, Union College

Abstract: Data have wound their way into every aspect of private-sector organizations, government agencies, academia, healthcare, and even our personal lives.  Further, data is the essential “raw material” for producing large language models and artificial intelligence (AI) systems. 

Unfortunately, even as the use of data has become more widespread, efforts to ensure it is of sufficiently high quality have lagged behind. Ten years ago, the cost of bad data to the US economy was estimated to be over $3 trillion per year. Today this cost is likely even higher. 

While there has been considerable research on data quality, the vast majority of the published literature has come from the computer science and information technology disciplines, rather than the statistics or data science disciplines. Further, the vast majority of this published literature focuses on the “data are right” problem, i.e., are the values accurate, timely, complete, not missing, and so on? There has been much less literature on the “right data” problem, i.e., are these data the best possible data that could be obtained to address a specific problem? Are they even minimally acceptable to address this specific problem? How would we know?

I have previously argued that given the indispensable role that data play in all forms of analytics, including AI, the “underwhelming” research on data quality has perhaps been the greatest missed opportunity in the history of the statistics discipline. Going forward, I argue that data quality still presents a golden opportunity to secure a bright future for statistics and data science.

 

Statistical Engineering: a roadmap for solving modern scientific & engineering problems

Simon Mak, Duke

Abstract: With rapid innovations in scientific computing and experimentation, engineers now have access to highly sophisticated experiments for tackling complex problems. Such sophistication, however, can come with high experimental costs and thus limited experiment runs. To effectively solve modern engineering problems, a statistical engineering approach is increasingly needed that integrates broad domain knowledge with statistical modeling and experimental design. In this talk, we explore the use of such approaches for tackling an application on rocket engine design via virtual experimentation. We develop a new boundary-informed Gaussian process surrogate model, which can incorporate known boundaries (from domain knowledge) of the simulated response surface. We further introduce an experimental design scheme that respects known boundaries, enabling the infusion of boundary information throughout the full model building process. Finally, we discuss ongoing work on discovering governing differential equations from experimental data.

 

Are “Fair Aware” Machine Learning Algorithms Really Fair?

Frederick W. Faltin, Virginia Tech

Abstract: A recurring concern with Machine Learning models is that their use will lead to unjust decisions or policies. This has led to proposals to assure “adequate” representation of select categories in the training data, and/or metrics for assessing the “bias” of modeling outcomes.

In some cases, model bias has been sufficiently flagrant that there is little doubt that some underlying problem exists. But in general, the operational definitions of “adequate” or “bias” being applied in ML contexts are anything but obvious. We as statisticians use the term “bias” in at least 2 ways: the “bias” of an estimator, or of a sample. Both of these are objective mathematical or statistical properties. But in society at large, the word “bias” almost invariably connotes moral or ethical shortcomings or misconduct. Which does the ML community mean? Are practitioners, or even researchers, aware of the distinction? Sample bias and moral/ethical bias may indeed intersect (generally with the former leading to the latter), but they’re not the same.

So what is the definition of model “bias” in “fair aware” usages, and what burden of proof must be met to demonstrate its presence? What if anything does the mathematics of “fair aware” algorithms give up in order to achieve their objective(s)? And might the answer to that question pose yet other concerns of fairness in sensitive applications, such as healthcare?

 

Bayesian Optimization for Hyperparameter Tuning in AI Models

Linfeng Lyu, Clemson University

Abstract: Hyperparameter tuning is a critical yet computationally expensive process in developing and deploying artificial intelligence (AI) models. The search space often involves a complex mixture of continuous, discrete, and categorical variables, making standard search methods inefficient. In this paper, we formulate hyperparameter tuning as a Bayesian optimization problem to efficiently navigate this mixed space under a strictly limited evaluation budget. We employ the Gaussian Process surrogate model combined with the Expected Improvement (EI) acquisition function to sequentially guide the search. We demonstrate the versatility and superiority of our proposed framework across diverse AI architectures, including Graph Neural Networks (GNNs), 3D Convolutional Neural Networks (CNNs), U-Net for image segmentation, and a GPT-based Large Language Model (LLM). Empirical results show that our approach consistently outperforms standard baseline methods, achieving faster convergence and improved model performance across various datasets.

 

Fitting an escalier to a curve

Sebastien Bossu, University of North Carolina at Charlotte

Abstract: We analyze the problem of fitting a fonction en escalier or multi-step function to a curve in L^2 Hilbert space, with a natural interpretation in terms of one-dimensional Binary Neural Networks. We propose a two-stage optimization approach whereby the step positions are initially fixed, corresponding to a classic linear least-squares problem with closed-form solution, and then are allowed to vary, leading to first-order conditions that can be solved recursively. We find that, subject to regularity conditions, the speed of convergence is linear as the number of steps n goes to infinity, and we develop a simple algorithm to recover the global optimum fit. Our numerical results based on a sweep search implementation show promising performance in terms of speed and accuracy.

Session 3 (1:30-3:00PM)

 

Credible Distributions of Overall Ranking of Entities

Abhyuday Mandal, University of Georgia

Abstract: Ranking, and inferences based on ranking of a set of entities,  are important problems in numerous contexts. This is especially true in small area statistics where there may be only a limited amount of directly observed data from each entity or small area, while precise and accurate estimates of best or worst performing entities are needed for fund allocation, planning and policymaking, stakeholder advocacy, evaluation of welfare programs,  and so on. However, ranks estimates constructed  exclusively on point estimates of parameters lack uncertainty quantification, and may lead to imbalances and inequities when these are based on small sample sizes. We propose novel Bayesian approaches to address this problem.  Our proposals result in partitions of the parameter space with  posterior distribution driven partial ordering of the sets in a partition. This in turn translates to a coherent probability mass function over ranks for every entity, and a coherent probability mass function over entities for every rank. Our Bayesian algorithms significantly outperform the state-of-the-art non-Bayesian alternatives, and are amenable to inclusion of covariates in the model as well as borrowing strengths across small areas.  We evaluate our proposed Bayesian algorithms in terms of accuracy and stability using  a number of applications and a simulation study. Additionally, we develop a novel theoretical framework for inference and ranking problems involving a triangular array of Fay-Herriot models and data, and provide probabilistic guarantees of performances of the proposed Bayesian ranking algorithms.

 

From Reviews to Insights: A Rank-Based Bayesian Model for Multimodal Aspect Sentiment Aggregation

Kaiwen Wang, Davidson College

Abstract: With the abundance of online reviews, it is of great interest for businesses to leverage the feedback from customers to inform business decisions. Specifically, when there is a consensus among online reviews on aspects that need improvement, it is paramount to accurately identify such consensus to inform actionable responses. While advanced large languages models (LLMs) are capable of extracting aspects and sentiments from individual reviews at ease, there is a growing need for an interpretable and scalable model that leverages both the upstream processing processing prowess of LLMs and the downstream Bayesian modeling with accurate parameter estimation. In this work, we propose a statistical pipeline based on a Bayesian rank aggregation model combining both latent variables and linear predictors for detecting the consensus among customer aspect rankings. Turbocharged with variational inference and multimodal upstream processing, the model efficiently performs consensus aggregation at a large scale. By applying our method to two large-scale datasets with hotel reviews, we showcase the practical applications of our model with important year-over-year insights.

 

AI-SALT: A Screening Accelerated Life Testing Framework for AI System Reliability

Simin Zheng, Virginia Tech

Abstract: Traditional accelerated life tests (ALTs) are a well-established method for assessing the reliability of physical products. However, as artificial intelligence (AI) technologies become increasingly widespread, and autonomous vehicles stand as a representative application, there is a growing need to extend ALT methodologies beyond physical components to evaluate AI systems. Whereas traditional ALTs are typically designed for a limited number of accelerated variables, AI ALT reliability tests may involve a large set of potential factors, making it infeasible to experiment with all of them simultaneously. Screening methods are widely used in the field of design of experiments, but their application to ALTs – particularly for AI systems – remains underexplored. This gap motivates the development of a comprehensive framework for screening-based ALT methods tailored to AI systems with responses in the form of recurrent event data. In this paper, we propose such a framework to identify and rank the most important variables, reduce a long list of candidates to a manageable subset, and reveal potential nonlinear relationships. Specifically, we adopt a flexible functional form for modeling covariates in the event intensity function and introduce a regularization term in the likelihood estimation to enable variable selection and facilitate statistical inference. Numerical optimization techniques are employed to obtain parameter estimates, ensuring both computational efficiency and reliable screening results. Finally, a physics-based simulation case study is presented to demonstrate the proposed screening ALT methodology for evaluating AI system reliability.

 

Panel – Journal Editors

Rong Pan, Arizona State University, Kamran Paynabar, Georgia Institute of Technology, and Bill Woodall, Virginia Tech

Abstract: Hear directly from the editors of Journal of Quality Technology, Quality Engineering, and Technometrics as they discuss the publication process from an editor’s perspective. Panelists will share insights on what makes a manuscript a good fit for their journal, common reasons for desk rejection, characteristics of successful submissions, and emerging trends in scholarly publishing. The session will include moderated discussion and audience questions, providing attendees with practical advice for navigating the publication process and communicating their work effectively.

 

CROSSVALI: A Method for Prompting AI for Data Analysis

Jennifer Van Mullekom and Anne Driscoll, Virginia Tech

Abstract: This talk is the evolution of the FTC 2025 talk, “Can ChatGPT Think Like a Quality Professional?”  Our data analysis prompting experience has prompted us (pun intended!) to create a framework for obtaining appropriate data analyses from artificial intelligence called CROSSVALI (Context, Refined Questions, Options, Specificity, Scrutiny, Verify & Validate, Ask Questions & Chain of Thought, Look Over, Interpret).  When communicating in a collaborative data analysis team, the team shares the burden of communication.  When collaborating with AI, the prompter takes on the majority of the burden of communication.  The CROSS portion of the framework is input focused to help the prompter create a well-defined, well-communicated prompt with the necessary elements to translate the domain question into a statistical one.  The VALI portion is human in the loop focused to ensure that we comply with ethical and responsible use of AI.  In this talk we will cover the details of the framework in the context of quality applications for both typical LLMs and agents using skills (e.g. Claude Code and associated skills).  At the end of the talk, attendees should be able to apply the framework and instruct teammates on the elements of the prompting framework in order to obtain more appropriate data analyses.

 

An Explainable Artificial Intelligence Framework for Child Trafficking Risk Assessment

Abiodun Olalere, North Carolina Agricultural and Technical State University

Abstract: Child trafficking in Sierra Leone’s Eastern Province remains a critical and multidimensional challenge. Findings from studies conducted between 2020 and 2024 reveal that many victims and survivors experienced multiple hazardous labor sectors and forms of exploitation, including force, fraud, and coercion. Children reported being subjected to a range of control techniques used by traffickers—such as physical violence, abusive relationships, social isolation, deprivation of basic needs, and manipulation of aspirations for a better future. Across the three districts studied, our data indicate shared patterns and characteristics between child trafficking (CT) and child labor (CL), reflected in overlapping risk profiles and cluster structures. For instance, children aged 12–17 were found to be at higher risk of trafficking compared with younger children aged 5–11.


This study employs an integrated explainable artificial intelligence (XAI) framework to examine and classify CT and CL typologies in relation to their underlying risk factors. Using latent class analysis alongside Classification and Regression Trees (CART), Random Forests (RF), CATBoost, XGBoost, and Artificial Neural Networks (ANN), we analyze 19 child trafficking indicators and one child labor indicator, as well as demographic (e.g., gender, age, family composition, education access), socio-economic (e.g., wealth index, labor contributions, displacement), and geographic (e.g., chiefdom-level) variables. The resulting models identify salient predictors of victimization risk and offer transparent, evidence-based insights into high‑vulnerability profiles and geographic risk areas. By applying XAI to original child trafficking data, this work advances the empirical and AI-driven insights needed to shape more precise, context-specific policies and interventions to reduce child trafficking and labor exploitation in Sierra Leone and other risk environments.

 

Space-Filling One-Factor-At-A-Time Designs

Wei-Yang Yu, Georgia Tech

Abstract: Space-filling designs are commonly used in deterministic computer experiments. However, they are ineffective for factor screening, which makes them inefficient when only a small subset of input factors is influential to the output. Recently developed screening designs, such as MOFAT designs, are effective at identifying important factors but lack space-filling properties, limiting their usefulness for surrogate modeling. In this article, we propose a new class of screening designs that improves the space-fillingness while retaining their screening capability. Through several numerical examples, we demonstrate that the proposed designs offer clear advantages over existing designs.

 Reception and Poster Session I (4:30-5:30PM)

 

Monitoring Supervised Learning Models with Statistical Control Charts

Leah Jones, Virginia Tech

Abstract: Machine learning (ML) is being integrated into a growing number of national security and defense systems, creating a need for continuous post-deployment monitoring to maintain reliable performance. Performance may deteriorate when the input data distribution changes (data drift) or when the underlying input-output relationship evolves (concept drift). This poster presents introductory considerations for applying control charting methods to monitor ML-enabled regression models and detect these forms of drift. We discuss where control charting concepts transfer directly and where they break down due to differences between traditional statistical and modern ML methods. We also summarize concept drift monitoring methods from prior literature and illustrate how they can be incorporated into control-chart frameworks. As a motivating case study, we focus on supervised learning with continuous outcomes.

 

Modeling Environmental Predictors of Grizzly Bear Scavenging in Yellowstone National Park

Becky Catlett, Virginia Tech

Abstract: In recent decades, grizzly bear (Ursus arctos) populations in the Greater Yellowstone Ecosystem (GYE) have grown after comprehensive recovery efforts. However, the effects of carcass availability, an important element of the grizzly bear’s food source, are underexamined. We analyze transect data spanning approximately two decades to investigate the relationship between the rate of carcass scavenging by grizzly bears in the presence of wolves and several environmental factors such as ungulate populations, elevation, winter severity, and grizzly bear density. This setting presents several statistical challenges, including possible spatial and temporal dependence and sparse observations of grizzly bear scavenging events. Utilizing a combination of logistic regression, comparative hypothesis testing, and spatiotemporal modeling techniques, we quantify how these variables influence scavenging behavior across the GYE. This analysis provides key statistical insights into carcass use across the interior of Yellowstone National Park, providing a quantitative foundation to evaluate and refine management strategies to support continued grizzly bear recovery.

 

A Two-Stage Structural Equation Modeling Framework to Address Spatial Confounding in Binary Response Models

Md. Hasibul Islam Jitu, North Carolina State University

Abstract: Confounding occurs when a variable affects both predictors and outcome. When that variable also follows a spatial structure that is not properly adjusted for, the resulting bias is known as spatial confounding. This creates a methodological challenge because unmeasured spatial dependence can influence associations between covariates and outcomes. In many public health and social science settings, an additional complication is that key determinants, such as socioeconomic status or household environment, are latent rather than directly observed. Structural Equation Modeling (SEM)-based logistic regression models can address latent constructs but do not account for spatial dependence. Logistic regression also ignores spatial dependence and does not incorporate latent constructs, while spatial models such as the Conditional Autoregressive Generalized Linear Mixed Model (CAR-GLMM) capture spatial dependence but do not include latent constructs. As a result, no framework exists for binary outcomes that jointly models latent constructs and accounts for spatial confounding.

Methods
This study proposes a two-stage Spatial SEM (SSEM) framework that integrates confirmatory factor analysis for latent constructs with a GLMM containing a Conditional Autoregressive random effect to represent spatial dependence, enabling joint modeling of latent constructs while accounting for spatial confounding in binary outcomes.

Results
Simulation experiments across varying spatial dependence, sample sizes, and latent-structure strengths show SSEM consistently outperforms logistic regression and SEM and achieves predictive performance comparable to, and often slightly better than, the CAR-GLMM, particularly under stronger spatial dependence and larger samples. Supplementary simulations show that SSEM achieves lower bias and RMSE than the SEM, indicating improved accuracy. In the real-data application using the Sample Vital Registration System, SSEM identifies key latent predictors of childhood deaths due to pneumonia and other respiratory diseases and achieves better predictive performance than the CAR-GLMM and non-spatial models. A de-spatialized predictor version is used to assess spatial confounding transmitted through indirect pathways. The stability of associations between the original and de-spatialized models suggests these indirect effects are weak, supporting robustness of the findings.

Discussion
This study introduces a framework that jointly models latent constructs and spatial effects, thereby accounting for spatial confounding, and achieves predictive performance and estimation accuracy comparable to or better than conventional approaches.

 

Estimation Stability Under Low Sample-Size-to-Parameter Conditions: Comparing Classical OLS and Regularized Bayesian Regression Estimators

Naa Adoley Acquaye, University of Denver

Abstract: Ordinary least squares (OLS) is widely used for estimating multiple regression models when classical assumptions are satisfied, especially when the number of observations exceeds the number of parameters and the design matrix is of full rank. When these conditions fail, however, OLS can become unstable and may no longer provide useful inference. This study compares the performance of the ordinary least squares method and the Bayesian method for multiple regression when the number of observations is less than or equal to the number of parameters. Using a real-world dataset, the sample size was reduced through random sampling so that the two methods could be examined under settings in which assumptions were both satisfied and violated. The comparison focused on whether the two methods produced similar estimates when n>kn > k and whether the Bayesian method remained more reliable when n≤kn≤k. Results showed that when n>kn > k and the model was full rank, the two methods produced similar estimates. In contrast, when n=kn = k, OLS produced unreliable parameter estimates with no standard errors, and when n<kn < k, OLS yielded unstable and limited results due to overfitting. The Bayesian method remained more robust across these conditions and produced more reliable estimates. These findings suggest that the Bayesian method is a stronger option for multiple regression in parameter-rich settings where classical methods break down.

 

Sequential Reliability Testing for LLM Code Generation

Vu Huynh, Virginia Tech

Abstract: We investigate whether sequential e-process certification can serve as a practical and statistically interpretable evaluation layer for large language model code generation. We develop a reproducible LeetCode-style benchmark in which each programming task is paired with generated assertion tests and a reference solution. For each task, an LLM produces candidate Python solutions through an iterative search phase using only coarse feedback from hidden tests. The selected solution is then evaluated in a separate certification phase using fresh, held-out tests. During both phases, we update a binomial likelihood-ratio e-process based on observed test failures, comparing the null hypothesis that a candidate’s failure probability exceeds an unacceptable threshold with an alternative representing a target level of reliability.


This framework separates two questions that are often treated as the same: whether feedback-guided search improves the code and whether the final candidate has accumulated enough statistical evidence to be considered reliable. The search phase records repair trajectories across attempts, while the certification phase freezes the selected candidate and evaluates it on independent test samples. The resulting outputs include failure-rate trajectories, empirical pass and fail rates, e-process capital, and stopping decisions. Preliminary experiments suggest that the framework can distinguish cases in which iterative search genuinely reduces failure rates from cases in which apparent progress does not lead to sufficient certification evidence. More broadly, the pipeline provides a transparent way to evaluate LLM-generated code under sequential uncertainty while connecting program testing with e-process-based and conformal-style ideas of validity.

 

Super Convergence of Instrument Assisted Subsampling Estimators

Jiaqi Liu, University of Connecticut

Abstract: Computational hurdles have become a significant bottleneck for massive data analysis in the era of big data. As data collection scales daily, computing demands frequently outpace available resources for full-dataset statistical analyses. A common workaround is uniform random subsampling to make computations viable; however, this approach is highly inefficient. While recent literature has proposed more efficient subsampling regimes to maximize estimation accuracy, improvements in efficiency are often marginal. Most existing estimators derived from a subsample of size r exhibit a standard convergence rate of r1/2 relative to the full-data estimator.


Although deterministic methods like Information-Based Optimal Subdata Selection (IBOSS) overcome these convergence limitations, they often fall short in practice due to stringent model and distributional assumptions. Real-world applications inherently involve varying degrees of model misspecification and irregularly distributed covariates, leaving stochastic subsampling as the more practical choice due to its robustness and ease of implementation.


To bridge this gap, we propose a novel stochastic subsampling estimation framework based on Poisson sampling, alongside two new super-convergent estimators. By utilizing an instrument, we uniquely leverage full-data information to unlock a super-convergence rate. We mathematically demonstrate that the proposed estimators are consistent with the full-data estimator, achieving accelerated convergence rates of r and r3/2.
 
Furthermore, this framework is highly flexible and practical. The proposed techniques can be applied to directly enhance existing estimators, making our method particularly appealing for modern applications where previously collected information must be leveraged efficiently as new data arrives continuously. Our estimators require only one-pass access to the full dataset and maintain computational complexities comparable to standard subsampling techniques. Numerical comparisons with existing methods confirm the best-in-class performance and practical utility of our proposed framework.

 

Estimating Decision Uncertainty from Preference Uncertainty: Application to Ground Vehicle Design

Chia-Ruei Liu, Clemson University

Abstract: Engineering design often involves selecting a solution from a Pareto set using a utility function derived from decision-makers’ preferences. However, these preferences are often uncertain, leading to variability in the resulting optimal design. This work proposes a probabilistic framework that models preference parameters as random variables and studies how this uncertainty propagates to decision outcomes. The resulting distribution over optimal designs reveals which regions of the Pareto front are most likely to be selected and provides a measure of recommendation stability. We further use global sensitivity analysis to identify key drivers of decision variability. A ground vehicle design case study demonstrates how the framework supports preference-aware decision making.

 

Respecting the Boundaries: Space-Filling Designs for Surrogate Modeling with Boundary Information

Yen-Chun Liu, Duke University

Abstract: Gaussian process (GP) surrogate models are widely used for emulating expensive computer simulations of complex physical phenomena. One challenge with fitting such surrogates is the costly nature of simulation runs, which greatly limits its training sample size given a fixed computational budget. Recent work has explored the integration of known boundary information for surrogate modeling, which shows promise for improving surrogate performance with limited data. Existing experimental designs, however, can be suboptimal for such boundary-integrated GPs, as they do not account for known boundary information. We thus propose a new class of space-filling designs, called boundary maximin designs, for effective GP surrogate modeling with boundary information. The proposed design jointly targets design space-fillingness and ensures design points are pushed away from known boundaries. We prove that boundary maximin designs have desirable information-theoretic properties, in that they are asymptotically D-optimal for existing boundary-integrated GPs. To account for effect sparsity, we further introduce a new boundary maximum projection design that jointly integrates boundary information and ensures good projective properties. Numerical experiments and a surrogate modeling application on particle collisions show the improved performance of the proposed boundary maximin designs compared to the state-of-the-art for boundary-integrated GPs.

 

csbewma: An R Package for Cumulative Standardized Binomial EWMA Control Charts

Faruk Muritala, Kennesaw State University

Abstract: The csbewma package implements the Cumulative Standardized Binomial Exponentially Weighted Moving Average (CSB‑EWMA) control chart for monitoring multiple independent binary streams. It provides exact, time‑varying mean and variance calculations derived analytically for the EWMA statistic, overcoming the limitations of asymptotic approximations in early monitoring phases. The package includes robust performance evaluation across four data‑generating distributions (normal, Laplace, uniform, exponential) and a post‑hoc diagnostic module that identifies out‑of‑control streams after a signal. The diagnostic module uses one‑sided binomial tests for each stream, followed by multiple testing corrections (Bonferroni, Holm, Benjamini‑Hochberg) and a simple max‑proportion rule. Extensive simulations confirm the chart’s robustness and the diagnostic methods’ trade‑offs between sensitivity and false discovery rate. The package is publicly available on CRAN. This poster will showcase the package’s features, the underlying theory, and practical examples for quality engineers and statisticians working with high‑dimensional binary data.

 

Conformal Prediction for Intrusion Detection

Chidiogo Onoh, University of Central Florida

Abstract: Standard intrusion detection systems issue point predictions without uncertainty metrics, hiding critical failures behind aggregate accuracy. This study evaluates whether conformal prediction can provide valid, finite-sample error guarantees for a logistic regression detector (N= 9,537, 44.7% attacks, AUC= 0.789). Although standard marginal conformal prediction meets its overall 90% coverage target (0.9009 across 200 splits), it silently under-covers the attack class in every split (mean 0.840). This stems from differential class difficulty because normal traffic is easier to separate so a single pooled threshold miscalibrates both classes. Applying class-conditional (Mondrian) conformal prediction resolves this failure, guaranteeing 90% coverage on attacks, reducing missed intrusions from 304 (35.6%) to 76 (8.9%), and adding only 2.3 percentage points to the referral rate. Referred sessions represent cases where model performance drops to baseline levels. Ultimately, this work demonstrates that default marginal guarantees fail silently on high-stakes classes and proves that Mondrian conformal prediction provides an effective, model-agnostic mechanism for security operations.

October 8th

Session 4 (8:30-9:15AM)

 

Probability of Detection: Evaluating the Reliability of Nondestructive Inspection Systems

Christine Knott, Air Force Research Laboratory

Abstract: The reliability of nondestructive inspection (NDI) systems is estimated using statistical methods, the most thorough of which is the Probability of Detection (POD) methodology. The Department of the Air Force uses periodic nondestructive re-inspection of critical structural components to maintain aircraft safety, and POD helps establish the length of inspection intervals. The established methods for POD will be provided, followed by a discussion of recent statistical research which extends and improves upon these methods.

 

Supersaturated Designs for Main Effects and Two‐Factor Interactions: Supersaturated Designs for Interactions

Maria Weese, Miami University

Abstract: Supersaturated designs (SSDs), which use fewer runs than factors, are employed in screening experiments to identify large main effects (MEs), even when the experimental setting may also involve factors that interact. This paper introduces new SSD design criteria that account for two‐factor interactions (2FIs) and compares them with designs from existing SSD approaches. Overall, we found that several established and new SSD criteria perform well for screening both MEs and 2FIs when analyzed by the hierarchical garrote, a recently proposed regularization method, while an existing foldover‐based method is promising when paired with the Lasso. When MEs are large and sparse, these designs can effectively identify them, except when the SSD is too small or excessively supersaturated. In addition, reasonably sized SSDs can reliably identify large strong heredity 2FIs; however, weak heredity interactions cannot reliably detected by any design considered.

 

Conjecturing-Based Discovery of Patterns in Data

David Edwards, The Citadel

Abstract: Automated scientific discovery seeks to identify meaningful relationships hidden within complex data sets and to generate hypotheses that can guide future investigation. We present a conjecturing-based framework for discovering interpretable nonlinear and Boolean relationships among features in observational data. The approach adapts the Dalmatian heuristic, originally developed for mathematical conjecture generation, to identify nonlinear bounds among numerical variables and logical relationships among categorical variables. Unlike traditional machine learning methods that emphasize prediction, the proposed framework is designed to reveal candidate relationships that may provide insight into the underlying structure of a system. Discovering patterns in data is a first step toward establishing causal relationships, which can form the basis for effective decision making.

Computational experiments demonstrate the ability of the framework to recover known nonlinear relationships and identify meaningful feature interactions from data. Comparisons with symbolic regression methods highlight the strengths and limitations of a conjecturing-based approach to pattern discovery. We also discuss practical challenges associated with evaluating and prioritizing large collections of conjectures, and present strategies for identifying those most likely to yield scientifically meaningful insight.

Session 5 (9:30-10:30AM)

 

Revisiting the 10p rule for Computer Experiments

Miles Woollacott, North Carolina State University

Abstract: For computer experiments with p factors, a popular rule-of-thumb followed by much of the existing literature is that n=10p runs are needed for reliable prediction. However, the papers from which this rule originates do not make such a strong claim. In this paper, we revisit the methodology in these papers and provide our own rule for necessary sample size. In particular, we find that the choice of design is important when taking into account parameter estimation. A surprising result is that randomly generated designs can have many advantages over traditional space filling designs. We also propose and justify a new sample size rule based on the effect hierarchy principle.

 

Profile Bayesian Optimization in Expensive Computer Experiments

Courtney Kyger, Virginia Tech

Abstract: We propose a novel Bayesian optimization (BO) procedure aimed at identifying the “profile optima” of a deterministic black-box computer simulation that has a single control parameter and multiple nuisance parameters. The profile optima capture the optimal response values as a function of the control parameter. Our objective is to identify them across the entire plausible range of the control parameter. Classic BO, which targets a single optimum over all parameters, does not explore the entire control parameter range. Instead, we develop a novel two-stage acquisition scheme to balance exploration across the control parameter and exploitation of the profile optima, leveraging deep and shallow Gaussian process surrogates to facilitate uncertainty quantification. We are motivated by a computer simulation of a diffuser in a rotating detonation combustion engine, which returns the energy lost through diffusion as a function of various design parameters. We aim to identify the lowest possible energy loss as a function of the diffuser’s length; understanding this relationship will enable well-informed design choices. Our “profile Bayesian optimization” procedure outperforms traditional BO and profile optimization methods on a variety of benchmarks and proves effective in our motivating application.

 

Optimal Sparse Projection Design for Systems with Treatment Cardinality Constraint

Kexin Xie, Virginia Tech

Abstract: Modern experiments often operate under treatment cardinality constraints, meaning that each treatment can include only a fixed number of factors. Such constraints arise naturally in engineering simulation, artificial intelligence system tuning, and large-scale system verification, where practical limits prevent the use of fully flexible designs. These settings require experimental designs that remain statistically efficient while respecting feasibility constraints. In this work, we study two-level designs under this setting, focusing on cases where every treatment contains the same number of active factors. Although these designs are closely related to balanced incomplete block designs, exact balance is not available for many practically relevant design sizes. This motivates the study of nearly balanced designs, which we show are optimal for the leading measures of aliasing in the generalized word-length pattern. We further show that strong projection performance in this setting depends on two simple structural properties: balanced use of individual factors and uniform co-occurrence of factor pairs. Based on this insight, we propose a new model-free design criterion that jointly penalizes imbalance in factor usage and irregularity in pairwise co-occurrence. We also establish close connections between this criterion and several classical design principles related to balance, projection quality, and Bayesian information efficiency. To construct high-quality designs, we develop a coordinate-exchange algorithm with efficient incremental updates, together with a simulation-based strategy for tuning criterion weights to the intended downstream task. Numerical studies show that the proposed approach performs favorably relative to existing alternatives across a wide range of problem sizes and constraint levels.

 

An Economical Approach to Design with Precision Criteria

Luke Hagar, The University of Queensland

Abstract: Estimation frameworks for statistical inference are preferred to hypothesis testing when quantifying uncertainty and precise estimation are more valuable than binary decisions about statistical significance. Study design for estimation-based investigations often uses precision criteria to select sample sizes that control the length of interval estimates with respect to a sampling distribution. We formally define a distribution that characterizes the probability of obtaining a sufficiently narrow interval estimate as a function of the sample size. This distribution can be used to determine the smallest sample size needed to ensure an interval estimate is sufficiently narrow. We prove that this distribution is approximately normal in large-sample settings for many data generation processes. However, this approximate normality may not hold for studies with moderate sample sizes, particularly when incorporating prior information or obtaining asymmetric interval estimates. Thus, we also propose an efficient simulation-based approach to approximate the distribution for the sample size by estimating the sampling distribution of interval estimate lengths at only two sample sizes. Our methodology provides a unified framework for design with precision criteria in Bayesian and frequentist settings. We illustrate the broad applicability of this framework with several examples.

 

Quantile-Based PCIs for Evaluating Two Processes

M Z Anis, Indian Statistical Institute

Abstract: One of the important uses of process capability indices (PCIs) is to quantify whether or not a (manufacturing) process meets the predetermined specification limits. Generally, the PCIs are estimated assuming the process is under statistical control, quality characteristics can be modelled by the normal distribution and the observations are independent.

However, many quality characteristics such as compositional data of some chemical processes (e.g. chromium plating), the lifetime of a component, cylindricity and taper in a manufacturing process follow a non-normal distribution. Measuring the performance of non-normal processes by using the classical elementary PCIs is inappropriate. There are two different approaches  to take care of such situations. The first one is to transform the non-normal data to a Gaussian distribution by using an appropriate transformation. The second approach is to modify the existing PCIs for the skewed distributions.  PCIs have been used to compare performance between processes and to select the best supplier among available suppliers. Two common methods are used to compare the PCIs of two suppliers:

(I)     100% test is performed to estimate the PCIs of each supplier separately; and subsequently the two suppliers are compared based on the true PCI values;

(II)     A sample is collected and statistical tests are performed to evaluate the two suppliers process capabilities.

Method (I) is time-consuming and expensive. Method (II) is also difficult to apply because a proper test static may not always be present for all PCIs. For these reasons, confidence intervals (CI) of the differences of the two PCIs are estimated to assess the process performances.

To the best of our knowledge, no such comparison has been made between the two processes by using quantile-based PCIs under general location-scale distributions. Furthermore, in most cases, the underlying distribution is assumed to be normal. In this work, we propose fiducial generalized confidence interval (FGCI) for the difference of CNpk and CNpm for testing the hypotheses H0 : CNpk1  = CNpk2 (H0 : CNpm1  = CNpm2 ) vs. H1 : CNpk1 ≠ CNpk2 (H1 : CNpm1 ≠ CNpm2 ) when the quality characteristics follow location-scale or log location-scale distribution. The proposed method can be applied to other two PCIs CNand CNpmk in a similar way.

We first introduce the concept of Fudicial Quantity (FQ) and discuss the problem of testing of hypothesis of comparison between two PCIs based on the fudicial generalized confidence interval of the difference of two PCIs. The procedure for estimating FQs of some important distributions is elaborated next. The method of estimating three important non-parametric bootstrap confidence interval for the difference of PCIs is subsequently discussed. Performance comparison is made based on Monte Carlo simulation. Two real life applications illustrate the utility of the proposed methods.

 

Predictive Ppk Calculations for Biologics and Vaccines using a Bayesian approach

Jianfang Hu, Pfizer

Abstract: In pharmaceutical manufacturing, especially biologics and vaccines manufacturing, emphasis on speedy process development can lead to inadequate process development, which often results in less robust commercial manufacturing process after launch. Process performance index (Ppk) is a statistical measurement of the ability of a process to produce output within specification limits over a period of time. In biopharmaceutical manufacturing, progression in process development is based on Critical Quality Attributes meeting their specification limits, lacking insight into the process robustness. Ppk is typically estimated after 15-30 commercial batches at which point it may be too late/too complex to make process adjustments to enhance robustness. The use of Bayesian statistics, prior knowledge, and input from Subject matter experts (SMEs) offers an opportunity to make predictions on process capability during the development cycle. Developing a standard methodology to assess long term process capability at various stages of development provides several benefits. We propose a Bayesian-based method to predict the performance of a manufacturing process at full manufacturing scale during the development and commercialization phase, before commercial data exists. Under Bayesian framework, limited development data for the process of interest at hand, data from similar products, general SME knowledge, and literature can be carefully formulated into informative priors. The implementation of the proposed approach is presented through two examples. To allow for continuous improvement during process development, we recommend to embed this approach of using predictive Ppk at pre-defined commercialization stage-gates, for example, at completion of process development, prior to and completion of PC, prior to technology transfer runs (Engineering/Process Performance Qualification, PPQ), and prior to commercial specification setting.

Session 6 (11:00AM-12:30PM)

 

XCT Systems Transfer Function Methodology Development Pathfinder: Linking System Performance Using Experimental Design

Mindy Hotchkiss, Enquery Research LLC

Abstract: X-ray Computed Tomography (XCT) systems are used to evaluate the quality of critical structural components for aerospace hardware at multiple locations. However, individual XCT systems are unique, with different capabilities and performance, so the “optimal” settings required to obtain a “good” quality scan vary from system-to-system, even for the same application. These optimal settings, as defined by a set of system-specific settings, are therefore not extendable, or transferable, to other systems. It was hypothesized that transfer functions relating different XCT systems could be developed by investigating and utilizing the relationships between specific system settings and various measures of performance. To evaluate the viability of this approach, performance outcomes for two XCT systems with different capabilities were explored by scanning the same test specimen under a range of conditions with a test matrix defined using structured Design of Experiments (DOE) methodology, to assess whether the “same” performance could be achieved, identify key contributors, and determine how best to mathematically model the identified relationships. The test specimen used in this study is a Ti6Al4V additively manufactured disk with multiple embedded feature types, such as cylindrical holes, cones, and grooves, and minimal bulk porosity.
XCT systems and its related processes are highly complex, comprising the following key stages: setup, data acquisition, image processing, and image reconstruction and analysis. Individual process variables, i.e., input factors, were known to have complex interactions with other input settings, both within and between process stages. However, the effect of changing input settings on specific performance measures, i.e., quality metrics, was not fully understood. XCT scan quality can be rated by several different quality metrics, each measuring a different aspect of scan quality. However, the same settings on an XCT system are not expected to be optimal for all quality metrics being evaluated. For this pathfinder study, the focus was limited to the XCT data acquisition process. Parameters involved in image reconstruction and analysis stages were defined, streamlined, and automated per Lawrence Livermore National Laboratory (LLNL) current best practice and locked down to reduce the scope and allow the experiment to focus on aspects of the system of primary interest, the system settings on the XCT system itself. 

The project goal and complex nature of XCT systems necessitated using structured experimental design, DOE, methods for the two disparate systems of interest studied. Two unique but similarly structured data acquisition experiments were designed using a more advanced DOE structure, in order to properly account for different points of system control, which creates a hierarchy that must be factored into both the test matrix design and the analysis to appropriately represent the system and model it correctly. The choices of input factors and associated levels for these factors were carefully considered, to best incorporate and account for the underlying physics of the XCT system and associated processes and interrelationships. The experimental design is currently being executed, as data is first collected using the XCT systems and then processed and analyzed to assess the values for each of the performance quality metrics. Preliminary statistical analysis results clearly show relationships between the quality metrics and the selected input factors, along with interactions between input factors. In addition, predictive models should provide additional insight into factor combinations to further optimize the XCT quality. A transfer function is also expected to be obtainable given sufficient overlap between system performance ranges. Success to date is considered directly attributable to the well-designed experiment that is the foundation for this extremely complex study.

 

Distributed Accelerated Failure-time Models for Big Data

Jared Clark, Virginia Tech

Abstract: Distributed computing enables the analysis of large scale reliability data that exceed the memory and computational limits of a single machine. The Backblaze hard drive failure dataset provides a motivating example, containing multi-year fleet level condition monitoring records with millions of drives and tens of billions of measurements. Such data are often too large to load or analyze centrally, making distributed inference necessary. We develop a distributed maximum likelihood framework for accelerated failure time models using the alternating direction method of multipliers. The data are partitioned across multiple worker nodes, each maintaining its own local copy of the model parameters. A consensus mechanism is used to ensure that all worker nodes agree on a common parameter estimate, and the algorithm iteratively coordinates local and global updates to achieve this agreement while controlling communication cost. This approach yields statistically valid inference for accelerated failure time models using data that reside on separate servers. Ultimately we hope to demonstrate the method on the Backblaze hard drive failure dataset and show that the distributed algorithm attains accuracy comparable to centralized estimation while achieving substantially improved scalability for modern big data reliability applications.

 

A Feature-Enhanced Capacity Prediction Method for Efficient Lithium-Ion Cell Grading under Multi-Temperature Conditions

Shih-Han Huang, Virginia Tech

Abstract: In battery pack assembly, cells are graded by measuring discharge capacity to enable consistent module performance. However, traditional grading tests often consider the capacity at a single temperature, which does not reflect the diverse environmental conditions for battery packs in real-world use. In this work, we present a data-driven grading method by utilizing features from the charging phase and the first half of the discharge phase to predict full discharge capacities under multiple temperature conditions. The proposed method enables early-stage prediction and supports grading based on the multi-temperature predictions and their uncertainties. Compared to conventional approaches, the proposed method reduces testing time and cost while improving grading robustness. The proposed method is validated across three major lithium-ion chemistries to demonstrate its merits. It also offers practical potential for quality control in battery pack manufacturing.

 

Panel – Beyond the Tenure Track: Careers in Industry and Applied Research

Ashley Childress, Eli Lilly, Todd Coffey, Pfizer, and Abby Nachtsheim, LANL

Abstract: For many students and early-career professionals, a tenure-track faculty position is a well-known and highly desirable career path. And yet, statisticians have opportunities to work in many different industries. Navigating the wide range of careers beyond academia can be challenging, though, without firsthand perspectives from those who have taken different paths.

Join us for an engaging panel discussion featuring three accomplished statisticians whose careers span multiple fields. Panelists will share their professional journeys, and attendees will gain insight into the diverse roles statisticians play outside traditional faculty appointments, the skills that enabled career transitions, and the rewards and challenges associated with different professional settings. A substantial portion of the session will be devoted to audience questions, providing participants with the opportunity to engage directly with panelists about career choices, professional development, work-life considerations, and strategies for building a successful career in statistics.

Whether you are a student exploring future possibilities, an academic considering alternative career paths, or an experienced statistician interested in learning about opportunities across sectors, this panel will offer valuable perspectives on the many ways statistical expertise can create impact beyond the tenure track.

 

Mechanism-Level Bayesian Inference for Latent Gaussian Models: A Scalable Framework for Spatial Flow

Michael R. Schwob, Virginia Tech

Abstract: Understanding directional flow across heterogeneous landscapes is a central objective in many scientific fields. Scalar potential surfaces provide mechanistic representations of transport (i.e., the flow of processes, such as gene flow or migration) with gradients encoding local flow direction and velocity. Bayesian dyadic models enable inference of such potential surfaces from pairwise observations, typically using latent Gaussian priors to capture spatial structure. However, standard Bayesian workflows target posterior inference for the entire latent field even when scientific conclusions depend only on low-dimensional functionals, such as local flow direction, boundary flux, or regional contrasts. This strategy becomes computationally prohibitive in large-scale applications and is misaligned with the scientific estimand. We develop a mechanism-level Bayesian computational framework that directly targets posterior distributions of scientifically meaningful functionals without reconstructing the full latent field, and we show that posterior uncertainty for these quantities can be expressed as quadratic forms containing the precision matrix and estimated efficiently using Krylov subspace methods. Beyond classical computational gains, the proposed approach aligns Bayesian inference with expectation-estimation approaches arising in quantum linear algebra. We demonstrate the approach in a simulation study and a landscape genomics case study, where we obtain accurate and efficient inference of directional gene flow and barrier structure.

 

Flexible Bayesian Variable Selection under Shape Constraints

Ayumi Mutoh, North Carolina State University

Abstract: A Gaussian process (GP) is a popular metamodel for non-parametric Bayesian regression, providing a flexible way to model unknown functions while also quantifying uncertainty. Complex systems often involve a large number of factors and identifying those with the most significant influence on the output is essential for efficient modeling. Variable selection in GPs typically operates through prior distributions on the mean and/or correlation parameters that induce sparsity with the posterior distributions. We introduce a new approach that can impose shape constraints, such as monotonicity, on a hierarchical, additive GP model. Elliptical slice sampling allows for rejection-free sampling of the shape-constrained GPs via transformations of latent, unconstrained GPs. A sparsity-inducing prior is assigned to each factor’s GP scale parameter to allow for fast Gibbs sampling. We focus our demonstration on the case of monotonic functions and also propose modifications to detect issues caused by model misspecification.

 

Validating Applications Based on Machine Learning Libraries Using Blocked Covering Arrays: A Case Study

Yeng Saanchi, JMP Statistical Discovery LLC & Sohyeon Kim, North Carolina State University

Abstract: The widespread adoption of machine learning (ML) libraries in decision-making systems has created an urgent need for principled and reliable validation methodologies. PyTorch, a widely used open-source library for deep learning in Python, is often integrated into larger software systems through application-specific wrappers. While the core library may be tested, these wrappers may introduce implementation errors that can adversely affect model performance and output reliability. Ensuring the correctness of such integrations is therefore critical.

In this work, we present a design of experiments-based approach for validating ML-enabled applications using blocked covering arrays, a combinatorial testing technique that enables efficient exploration of input parameter spaces. A strength-t blocked covering array guarantees coverage of all t-way combinations of factors while also ensuring complete coverage of (t–1)-way combinations excluding the designated blocking factor. We treat hardware architecture as the blocking factor, motivated by its well-documented impact on ML model behavior and reproducibility.

We demonstrate the effectiveness of this approach through a case study involving regression and image classification tasks implemented via the JMP Torch Deep Learning Add-in. Our results suggest that blocked covering arrays provide a scalable and systematic framework for detecting inconsistencies and validating ML application wrappers under varying environmental conditions.

Session 7 (2-3:30PM)

 

AI-Augmented Use Risk Management for Medical Devices Under ISO 14971

Venudhar Hajari, Smith & Nephew plc

Abstract: Use related risk management remains one of the weakest areas in medical device development. Although ISO 14971 and IEC 62366 require systematic identification of use errors and hazardous situations, current practice is manual, fragmented, and often performed late in the lifecycle. As a result, critical use hazards are frequently discovered during human factors validation, when design changes are costly and constrained.
 
This paper presents an AI-augmented workflow that embeds hazard reasoning directly into the Use Risk Analysis process. Natural language models parse user tasks, infer plausible use errors, generate hazardous situations, and propose candidate risk controls across software, UI, hardware, and labeling. The system maintains live traceability between tasks, hazards, harms, and mitigations, enabling continuous risk evolution as designs change.
 
The approach is demonstrated using a real RF surgical system. From a single procedural task, the model identifies error modes such as unintended activation or prolonged energy delivery and maps them to patient harms and control strategies. This exposes failure chains that are routinely missed in traditional URRA and surfaces them early, before design freeze.
 
Results show earlier discovery of high-severity use hazards, stronger traceability between HF tasks and ISO 14971 artifacts, and risk-driven feedback during concept and architecture phases. This is not AI for documentation. It is a shift from reactive compliance to proactive safety engineering, providing a practical, standard-aligned evolution of use risk management.

 

Dark Uncertainty Imperils Hybrid Metrology for Semiconductors

Adam Pintar, National Institute of Standards and Technology

Abstract: Semiconductor manufacturing demands difficult measurements and audacious analyses. As device dimensions shrink, so too must the corresponding measurement uncertainties. For example, an influential technology roadmap has targeted 95 % coverage intervals of less ± 0.2 nm for device linewidths by 2028, championing hybrid metrology as the solution. In essence, hybrid metrology combines results from a few different methods with complementary principles of measurement using a statistical model. Several variations on this theme have emerged since its introduction to semiconductor manufacturing more than a decade ago. All approaches to hybrid metrology share two characteristics of requiring consistent results for combination and compelling a combined uncertainty that is less than the uncertainty of each individual method.

When individual methods produce consistent results―such as, on the occasion of a reliable inter-method comparison that demonstrates measurement results deviating by less than their uncertainty estimates―a reduction of combined uncertainty may be justifiable. However, in the broader context of difficult measurements, inconsistency is more common than consistency among the results of different methods. The issue is so prevalent that inconsistency between methods lacking an identifiable cause is known as dark uncertainty, a term that acknowledges both its invisibility when considering results from only a single method, as well as its prevalence. A rigorous justification of for hybrid metrology presents many challenges, including different measurands, modeling errors, and non-uniform samples.

In this study, we consider four main topics to assess the prospects for saving hybrid metrology from dark uncertainty. First, we explore the latent connections between hybrid metrology, inter-laboratory studies, and meta-analyses. Second, we illustrate potential causes of dark uncertainty in semiconductor manufacturing measurements through examples of linewidth and film thickness. Third, we consider several analyses to combine results from a few different methods. We test analyses that fail to acknowledge dark uncertainty, resulting in a combined uncertainty that is smaller but misleading, as well as analyses that acknowledge dark uncertainty, typically resulting in a combined uncertainty that is larger but more realistic. Fourth, we discuss strategies for reducing dark uncertainty by improving the scientific discourse between metrologists applying different methods and statisticians combining their results.

 

Monitoring Parametric, Nonparametric, and Semiparametric Linear Regression Models using a Multivariate CUSUM Bayesian Control Chart

Abdel-Salam G. Abdel-Salam, Qatar University

Abstract: Effective decision-making during epidemics requires not only forecasting case counts, but also continuous monitoring of the underlying epidemiological dynamics. This presentation introduces a multivariate statistical process monitoring framework for tracking time-varying parameters of compartmental epidemic models, with a focus on the Susceptible-Exposed-Infected-Recovered-Death-Vaccinated (SEIRDV) model. The approach integrates Bayesian sequential parameter estimation with Multivariate Exponentially Weighted Moving Average (MEWMA) and Multivariate Cumulative Sum (MCUSUM) control charts to enable real-time detection of both gradual and abrupt changes in disease transmission, recovery, mortality, and vaccination effects.

Using COVID-19 data from the State of Qatar as a motivating application, the talk demonstrates that monitoring estimated model parameters rather than observed case counts alone provides early warning signals for public health interventions, policy changes, and emerging variants. Simulation studies compare the sensitivity and robustness of MEWMA and MCUSUM charts across varying magnitudes of parameter shifts, highlighting trade-offs between rapid detection and false-alarm control.
 
The proposed framework bridges epidemiological modeling, Bayesian computation, and multivariate quality control, offering a general methodology applicable to a wide range of infectious diseases and dynamic systems. The talk emphasizes practical implementation, interpretability, and the role of statistical monitoring in supporting timely, data-driven public health decisions. 

 

Panel – From Bench to Batch: Nonclinical Statistics Case Studies from the Pharmaceutical Industry

Briana Russo, Merck, Greg Steeno, Pfizer, and Kade Young, Eli Lilly

Abstract: This panel features three experienced nonclinical statisticians who will showcase real-world applications of statistics across the drug development lifecycle, from early laboratory experimentation to commercial manufacturing. Through a series of practical case studies, panelists will demonstrate how statistical methods are used to address challenges in drug discovery and development. Attendees will gain insight into the breadth of problems encountered in pharmaceutical development and the impact statisticians can have in accelerating scientific understanding, improving process robustness, and supporting data-driven decision making. This panel offers an exciting opportunity to explore a vibrant and growing area of applied statistics while connecting with professionals working at the intersection of science, engineering, and data.

 

An Interpretable Generative Framework for Statistical Process Control in High-Dimensional Settings

Waldyn Martinez, Miami University

Abstract: Modern quality control settings increasingly involve high-dimensional, autocorrelated, and nonlinear process data for which classical multivariate control charts may be inadequate. Traditional approaches such as Hotelling’s T^2, MEWMA, and MCUSUM are often built on restrictive distributional assumptions and may struggle when the in-control structure is complex, time-dependent, or contaminated by atypical observations. To address these limitations, we propose ReGEN-TAD (Refined Generative Ensemble for Temporal Anomaly Detection), a deep learning-based monitoring framework designed for statistical process control in both Phase I and Phase II applications.
 
In Phase I, ReGEN-TAD is used to analyze historical process data, identify special-cause observations, and refine the reference set used for calibration. The method combines temporal representation learning with a refinement mechanism that filters anomalous observations from the baseline sample, thereby improving the construction of in-control limits. In Phase II, the calibrated model is deployed sequentially as a control-charting tool for prospective monitoring, producing anomaly scores that signal departures from the in-control state in real time. This yields a unified framework in which Phase I data cleansing and Phase II surveillance are handled by the same underlying methodology.
 
The proposed approach is especially suited to modern manufacturing and process environments where shifts may be subtle, persistent, nonlinear, or embedded in cross-sectional dependence. By learning latent temporal structure directly from the data, ReGEN-TAD avoids strong parametric assumptions while retaining the control-chart objective of distinguishing common-cause from special-cause variation. Results from simulated and benchmark process scenarios indicate that the method achieves strong detection capability with competitive false alarm control, particularly under challenging departures from normality and independence. These findings suggest that ReGEN-TAD provides a flexible and effective alternative for next-generation multivariate SPC and control chart design.

 

Beyond Prediction: Monitoring Dynamic Risk Using Machine Learning and SPC

Mohammad Abdullah, The Ohio State University

Abstract: In many modern applications, predictive models are used to estimate the probability of future events for individual entities over time, including domains such as cybersecurity, manufacturing, healthcare, and finance. However, most existing approaches emphasize prediction alone and lack a structured mechanism for monitoring how these probabilities evolve relative to similar entities. As a result, it becomes difficult to distinguish between normal variation and meaningful deviations that require intervention. This work introduces a new statistical monitoring framework, the Probability-based Individual Longitudinal Group-Referenced Risk Chart for Interpreting and Monitoring (PILGRIM). The proposed approach integrates machine learning with Statistical Process Control (SPC) by monitoring predicted probabilities rather than raw observations. Each entity is tracked longitudinally and evaluated against a reference group of similar entities at the same time point. A key concept is Conditional Statistical Control, where control is defined relative to the current group behavior instead of a fixed historical baseline, allowing the system to evolve while still identifying abnormal entity-level patterns. Control limits are constructed on the model scale and mapped to the probability scale, with theoretical guarantees for false alarm control based on Average Run Length (ARL). Overall, this research presents a novel and practical framework that bridges predictive modeling and statistical monitoring, enabling interpretable and reliable tracking of risk in dynamic, evolving environments.

 

New Perspectives of Respondent-Driven Sampling with Application to Sex-Trafficked Population in Senegal

Hui Yi, North Carolina Agricultural and Technical State University

Abstract: Respondent-Driven Sampling (RDS) is important for estimating prevalence with hidden and sensitive populations under restrictive fieldwork conditions. In this work, we address critical methodological challenges, including unconverged data and structural network fragmentation from a sample of 561 women aged 18–30 engaged in commercial sex in the Kédougou (urban) and Saraya (rural) departments of Senegal. Prevalence was estimated using different approaches, including Salganik–Heckathorn, Volz–Heckathorn, and Homophily Configuration Graph estimators. Further diagnostics are conducted using bootstrap-based uncertainty quantification and rigorous sensitivity analyses, including Leave-One-Seed-Out diagnostics. The Leave-One-Seed-Out analyses revealed substantial seed-level influence in the urban sample, where the exclusion of specific deep recruitment chains materially shifted prevalence estimates. It indicates that inference can remain tethered to initial starting conditions even when wave-based convergence appears stable. Our findings suggest that in hidden settings, data-structural limitations often supersede estimator selection. We conclude that for RDS to yield credible prevalence estimates in high-stakes research, one should incorporate seeding-effect diagnostics and wave-specific checks into the existing framework to ensure findings to be robust and accurate.

Reception and Poster Session II (3:30-4:30pm)

 

Sparse Bayesian Optimization on the Simplex via Manifold-Aware Kernels

Vivek Singh, Duke University

Abstract: Many optimization problems in mixture design, such as choosing compositions of materials or chemicals live on the probability simplex rather than in an unconstrained Euclidean space. In practice, only a small subset of the components meaningfully affects the response, while the rest add noise. Existing high-dimensional Bayesian optimization methods typically handle this sparsity through per-dimension length-scale priors, an idea built for Euclidean space that is not well posed on the simplex, since the constraint that components sum to one removes any independent per-coordinate direction to attach a length-scale to. We developed a Bayesian optimization method for the simplex that separates the question of which components matter from the question of how to measure distance between compositions. The approach combines a variable selection mechanism with a kernel built on the intrinsic geometry of the simplex, so that distances remain meaningful even near the boundary, where sparse, near-pure compositions often live. The method offers a principled path to sparse optimization on constrained, high-dimensional composition spaces without importing a Euclidean sparsity mechanism that doesn’t transfer to the simplex. Early results on synthetic benchmarks show meaningful gains over standard baselines, particularly in settings where the optimum lies near the boundary of the composition space, a common feature of real mixture design problems.

 

Exploring Intrinsic Gaussian Processes as Priors for Estimation Problems

Colby Stakun-Pickering, Virginia Tech

Abstract: Stationary Gaussian processes (GPs) have long been a staple as priors for spatial fields to be inferred in a wide variety of estimation problems. GPs are proper distributions, defined by their mean and covariance functions. GPs have also proven remarkably extendable, leading to innovations such as deep GPs and Vecchia-based approximations for large systems. While the standard, stationary GP has an established track record, prominent researchers from the past (e.g., Matheron in the ’70s, Besag in the ’90s) have advocated intrinsic GPs in their stead. These intrinsic formulations are improper and add some advantages, along with theoretical and computational hurdles. This poster will (re)introduce intrinsic GPs and explore how they might be used in Bayesian estimation problems and how the resulting posteriors compare to stationary GP-based formulations. This is joint work with a number of colleagues at Virginia Tech.

 

A Comparative Simulation Study of Signal/Response and Hit/Miss Methodologies for POD Analysis

Kaelan Swallow, Strategic Ohio Council for Higher Education

Abstract: Reliability assessments for nondestructive evaluation (NDE) systems are approximated using statistical methodology. The more robust method being that of Probability of Detection (POD). To maintain aircraft safety the Department of the Air Force (DAF) uses repeated nondestructive re-inspection of critical structures and POD aids in establishing inspection intervals. Various established methodologies for POD will be given, followed by discussion on design of statistical experiments and introduction of early results. This study focuses on how results from using Signal/Response (SR) and Hit/Miss (HM) POD methodologies differ.

 

A Geometric Approach to High-Dimensional Statistical Process Control via Manifold Fitting

Ismail Burak Tas, The Pennsylvania State University

Abstract: Multivariate Statistical Process Control (SPC) methods based on principal components or other latent space techniques rely on linear dimensionality reduction, which can be insufficient when monitoring factory-wide processes or large sets of manufacturing steps operating under hundreds of feedback loops. The normality assumption underlying most multivariate SPC methods becomes increasingly indefensible as dimensionality grows and the difficulty of the SPC problem further increases in the presence of autocorrelation.

We present a new approach for statistical control of high-dimensional, dynamical processes under a (nonlinear) manifold assumption. Most  manifold learning algorithms, however, cannot be used for on-line monitoring because they lack an explicit out-of-sample extension that would permit sequential data processing. Instead, the proposed method works directly in the ambient phase space of the dynamical process, and takes a geometrical point of view: a manifold is fit pointwise to the in-control (Phase I) data, without reducing dimension or filtering autocorrelation, and the deviations from the fitted manifold are computed. These deviations are monitored by a novel univariate, distribution-free control chart.

It is shown how the new Deviations-from-manifold (DFM) SPC method has theoretical detectability properties,  how it has a  controllable Type I error probability (or in-control average run length) and  can operate in the high-dimensional regime where ambient dimension exceeds the number of Phase I in-control observations. Numerical experiments on synthetic processes and on a large chemical process simulator are used to compare the performance of the proposed DFM approach with manifold-learning and autocorrelation-filtering monitoring alternatives that use comparable nonparametric detection approaches.

 

Multiclass Calibration Assessment and Recalibration of Probability Predictions via the Linear Log Odds Calibration Function

Amy Vennos, Virginia Tech

Abstract: Machine-generated probability predictions are essential in modern classification tasks such as image classification. A model is well calibrated when its predicted probabilities correspond to observed event frequencies. Despite the need for multicategory recalibration methods, existing  methods are limited to (i) comparing calibration between two or more models rather than directly assessing the calibration of a single model, (ii) requiring under-the-hood model access, e.g., accessing logit-scale predictions within the layers of a neural network, and (iii) providing output which is difficult for human analysts to understand. To overcome (i)-(iii), we propose Multicategory Linear Log Odds (MCLLO) recalibration, which (i) includes a likelihood ratio hypothesis test to assess calibration, (ii) does not require under-the-hood access to models and is thus applicable on a wide range of classification problems, and (iii) can be easily interpreted. We demonstrate the effectiveness of the MCLLO method through simulations and three real-world case studies involving image classification via convolutional neural network, obesity analysis via random forest, and ecology via regression modeling. We compare MCLLO to four comparator recalibration techniques utilizing both our hypothesis test and the existing calibration metric Expected Calibration Error to show that our method works well alone and in concert with other methods.

 

Sequential IMSE Design for Manifold-Regularized Kernel Logistic Regression

Xiaotian Wang, University of Georgia

Abstract: Efficient label acquisition is a central challenge in high-dimensional classification, particularly when unlabeled observations are abundant but class labels are costly to obtain. We develop a sequential experimental-design framework for binary classification under the assumption that the covariates concentrate near a lower-dimensional manifold. Class probabilities are modeled using manifold-regularized kernel logistic regression, which combines a reproducing kernel Hilbert space norm penalty with a graph-Laplacian penalty constructed from labeled and unlabeled covariates. At each stage, the next observation to label is selected by minimizing an approximation to the weighted integrated mean squared error of the estimated class probabilities over a prespecified prediction domain. The criterion incorporates both the shrinkage bias and the estimation variance induced by the structured quadratic penalty. By defining the candidate set separately from the prediction domain and its weighting measure, the framework directly targets predictive accuracy over scientifically relevant regions rather than only over the observed unlabeled pool. The ambient and intrinsic regularization parameters are re-estimated from the accumulating labeled data at each stage, allowing the balance between RKHS smoothness and manifold smoothness to adapt throughout the sequential design. The resulting method extends experimental-design-based active learning for penalized logistic regression to manifold-aware probability estimation and provides a principled strategy for allocating a limited labeling budget.

 

Sequentially Weighted One-Factor-at-a-Time Designs for Safe Online Controlled Experiments

Rachel Winpisinger, North Carolina State University

Abstract: Online controlled experiments (OCEs) are sequential experiments widely used in industry to evaluate new product features and optimize digital services. OCE practitioners are often required to balance learning about multiple experimental factors while minimizing user exposure to poorly performing treatments. Multi-armed bandits (MABs), a common class of sequential optimization methods, are designed to maximize cumulative reward and therefore often reduce exploration before sufficient information has been collected to adequately characterize the treatment space for the objectives of these studies. To address these limitations, Haizler and Steinberg (2021) proposed Fractional Factorial Design + Thompson Sampling (FFD+TS). This method incorporates factorial structure into sequential experimentation, enabling more informed exploration while balancing exploration and exploitation. We identify limitations of this approach and propose a framework for safe sequential factorial experimentation. The framework combines baseline parametrization and adaptive allocation to guide exploration while maintaining unbiased estimation of treatment effects using only a limited subset of treatment combinations. We implement this framework through Sequentially Weighted One-Factor-at-a-Time (SW-OFAT), a sequential heuristic that adaptively updates both treatment allocation and the experimental baseline as information is collected. Simulation studies compare SW-OFAT with FFD+TS and three widely used MAB methods, demonstrating competitive or improved regret and treatment selection performance while providing a safer and more interpretable approach to sequential factorial experimentation.

 

Monitoring 75,000 Calorimeter Channels: Coupled, Anytime-Valid Sparse Changepoint Detection

Patrick Woitschig, Duke University

Abstract: The electromagnetic calorimeter in the CMS experiment at the Large Hadron Collider contains approximately 75,000 crystal cells whose calibration and detector response require continuous monitoring. Changes caused by radiation damage, electronics faults, calibration shifts, or changes in operating conditions may affect one cell or many cells simultaneously. In addition, several fault modes are possible, and each may produce a different, often non-Gaussian post-change distribution. The affected cells, the time of change, and the underlying fault mechanism are therefore all unknown. This creates an extreme-scale statistical process-monitoring problem in which repeated testing can generate an unacceptable false-alarm burden.

We present a scalable framework for high-dimensional, coupled changepoint detection that combines frequentist false-alarm guarantees with prior-weighted evidence aggregation. For each cell, we construct nonnegative evidence scores called e-values. Different evidence models can represent different candidate fault modes, while remaining valid across a family of no-change models even under continuous monitoring. This cellwise construction accommodates non-Gaussian observations through application-specific evidence models. A coupled participation prior then combines the evidence over possible affected subsets. It allows multiple cells or regions to change simultaneously and can encode expected sparsity and physical relationships, including detector geometry, shared readout components, and scientifically plausible failure patterns.

The aggregated evidence drives e-value versions of the Shiryaev–Roberts and CUSUM procedures. The resulting detectors provide finite-sample average-run-length control under composite no-change models. Direct anytime-valid thresholding can also be used to control the probability of a false alarm. The method remains computationally tractable at calorimeter scale because it combines evidence across cells without explicitly examining every possible group of affected cells. An adaptive beta-Bernoulli construction further accommodates an unknown fraction of affected cells.

The participation prior has an operational frequentist interpretation in addition to its Bayesian-style modeling role. Its negative log weight represents the unavoidable allocation cost of competing with an oracle that knew the affected cells in advance. This provides a principled way to balance sensitivity to isolated faults against sensitivity to spatially extended or detector-wide changes.
A calorimeter-monitoring application demonstrates the complete workflow: constructing valid, fault-specific evidence at the cell level, coupling evidence across physically related cells, issuing an online alarm with controlled false-alarm behavior, and localizing the cells most responsible for the alarm. The modular separation of fault modeling, participation modeling, online stopping, and post-alarm diagnosis makes the framework applicable to large detector systems and other high-dimensional sensor networks.