Date Created

5-2026

Embargo Date

5-2028

Document Type

Dissertation

Degree Name

Doctor of Philosophy

Department

College of Education and Behavioral Sciences, Applied Statistics and Research Methods, ASRM Student Work

First Advisor

Tsai, Chia-Lin

First Committee Member

Merchant, William

Second Committee Member

Yu, Han

Third Committee Member

Paek, Sue Hyeon

Abstract

This study focuses on evaluating data treatment methods such as truncation and top-coding in the context of a difference-in-differences analysis of healthcare expenditure data, in which a two-part model is fit with a generalized linear model using gamma distribution and log link in the second part of the model. It starts with a deep dive into the nature of healthcare cost data and the source of these data hallmarks, particularly why these data (and those of other fields) give rise to substantial mass at zero and exhibit long, right tails, and the theoretical reasons why these values should not be excluded from the dataset. The majority of this paper focuses on the handling of the long, right tails (or extreme values) often present in these datasets. Chapter II continues to explore methods of identifying and handling right-skew. First, it describes the issues of mixed distributions, heteroskedasticity, and severe right skew in more detail. Then it provides an overview of the various techniques (data transformation, data treatment, and modeling) that have been suggested and applied in the literature. The study concludes with a thorough analysis of five data treatment methods’ impact on bias, precision, and confidence interval coverage of the interaction term treatment group × prepost as well as findings related to average treatment effect on the treated (ATT) estimates. Overall, the study found that no single method dominated across all data scenarios. Standard top-coding did, however, offer the most consistent performance across bias, standard error precision, confidence interval coverage, and ATT, particularly among scenarios that preserve the full population estimand; illusory advantages offered by truncation were dependent on condition and accompanied miscalibrated uncertainty. As the performance of each data treatment method varied systematically across estimands and conditionally upon data scenarios, researchers should select a data treatment strategy based on estimands most central to their research question and overall study design.

Abstract Format

html

Language

English

Extent

239 pages

Rights Statement

Copyright is held by the author.

Digital Origin

Born digital

Available for download on Monday, May 01, 2028

Share

COinS