table1() builds the descriptive "Table 1" of a study in one call: numeric
variables as mean (SD) or median (IQR), categorical variables as counts and
percentages. Add by to get one column per group, optionally with
p-values and crude or adjusted effect measures (PR, RR, or OR). For a single
cross-tabulation use tb(); for full regression output use regtab().
Usage
table1(
data,
vars,
by = NULL,
...,
flags = character(),
measure = NULL,
adjust = NULL,
test = FALSE,
summary = "auto",
overall = TRUE,
missing = TRUE,
denominator = c("available", "complete"),
na_model = c("drop", "fail", "explicit"),
ref = NULL,
var.type = NULL,
labels = NULL,
style = "default",
d = 1,
conf.level = 0.95,
design = NULL
)Arguments
- data
A data frame.
- vars
Variables to describe: bare names in
c(), a character vector, or a tidyselect expression such asstarts_with("cm_").- by
Optional grouping variable: a bare name or string. Produces one column per group.
- ...
Terse flags such as
poror, written afterby; see Terse flags. Only flags are allowed here.- flags
Character vector of terse flags, e.g.
c("p", "or"). The programmatic form of the bare flags in....- measure
String. Effect measure to add:
"PR","RR", or"OR". Requires a binaryby, whose last level is taken as the event.- adjust
Optional covariates for an adjusted effect-measure column, given like
vars. Requiresmeasure.- test
Logical or string.
TRUEadds a p-value column with an automatically chosen test. A string forces one test:"chisq","fisher","mcnemar", or"trend"for categorical variables, and"t","wilcoxon","anova", or"kruskal"for numeric ones.- summary
String. Summary for numeric variables:
"auto"chooses between"mean"(mean and SD) and"median"(median and IQR) based on sample size and skewness. A named list sets it per variable, e.g.list(age = "mean", bmi = "median").- overall
Logical. If
TRUE, add an "Overall" column for the whole cohort.- missing
Logical. If
TRUE, add a "Missing" count row under each variable that has missing values. This does not change any other number.- denominator
String. Which rows the percentages and summaries use:
"available"uses every non-missing value of each variable;"complete"uses only rows complete on all ofvarsandby.- na_model
String. How adjusted models handle missing values:
"drop"removes incomplete rows,"fail"stops with an error, and"explicit"treats missing categorical predictors as a"(Missing)"level.- ref
Reference level for effect measures: one level for all variables, or a named list per variable. If
NULL, each variable's first level is used.- var.type
Force variable types:
"continuous"or"categorical", either one string for all variables or a named vector such asc(score = "continuous"). IfNULL, types are detected automatically (see Statistical methods).- labels
Named character vector of display labels, e.g.
c(smoking = "Smoking status"). Variable labels already stored indataare used by default.- style
Display style: a journal preset name (see
list_journals()), ajournal_style()object,"n_pct"or"pct_n", a template using{n}and{p}, or a named list per variable.- d
Integer. Decimal places for percentages and continuous summaries.
- conf.level
Number between 0 and 1. Confidence level for effect measure intervals.
- design
String. Study design used to choose the effect measure when
measureis not set:"cross_sectional"gives PR,"cohort"gives RR, and"case_control"gives OR.
Value
A simtab_result of class simtab_table1. Print it to see the
formatted table, convert it with as.data.frame(), or save it with
export_docx(), export_pptx(), or export_xlsx(). Unrounded results
are stored in $data.
Details
Statistical methods
Crude effect measures for categorical variables use the same estimators as
tb(): the Katz log interval for prevalence and risk ratios (Katz et al.,
1978) and the Woolf logit interval for odds ratios (Woolf, 1955). For
numeric variables, and for every adjusted estimate, a separate generalized
linear model is fitted for each described variable, with the adjust
covariates added. Odds ratios come from logistic regression. Prevalence and
risk ratios come from log-binomial regression, falling back to modified
Poisson regression with robust standard errors when that model does not
converge (Barros & Hirakata, 2003; Zou, 2004). PR and RR are computed identically; only the label
differs.
With test = TRUE, 2x2 tables use the N-1 chi-squared test (Campbell,
2007) and larger tables the Pearson chi-squared test, without continuity
correction; Fisher's exact test is chosen only when an expected count is
below 1. Numeric variables are compared with Welch's t-test or ANOVA when
summarised by the mean, and with the Wilcoxon or Kruskal-Wallis test when
summarised by the median.
Numeric variables are treated as continuous unless they look like coded
categories: exactly two distinct whole numbers, or at most seven distinct
whole numbers with at least 20 observations. Use var.type to override.
Missing data
By default each variable is described using its own non-missing values,
and a "Missing" row reports how many are absent (STROBE item 14). Tests are
always computed on observed values. denominator = "complete" restricts
the whole table to complete cases, so the displayed numbers can change.
na_model controls only the adjusted models.
Modifying the result
Multiplicity adjustment, paired tests, and standardized mean differences
(Austin, 2009) are set afterwards with test(), e.g.
table1(df, vars, by = group, test = TRUE) |> test(p.adjust = "holm", smd = TRUE).
With paired = TRUE, two equal-sized groups are treated as matched pairs in
row order. Use add_effect() to add or replace the effect-measure columns
and fmt() to change decimals.
Terse flags
Flags may be supplied as bare words in ... after by, or programmatically with flags =. Anything else in ... is treated as a typo.
colColumn percentages (the default).
row,cell,percNot yet supported by
table1().pr,rr,orAdd the corresponding crude effect measure.
pAdd a p-value from an automatically chosen test, like
test = TRUE.missShow missing values.
The old flags rp and m still work as aliases for pr and miss but are deprecated.
References
Austin, P. C. (2009). Balance diagnostics for comparing the distribution of baseline covariates between treatment groups in propensity-score matched samples. Statistics in Medicine, 28(25), 3083–3107. doi:10.1002/sim.3697 .
Barros, A. J. D., & Hirakata, V. N. (2003). Alternatives for logistic regression in cross-sectional studies: an empirical comparison of models that directly estimate the prevalence ratio. BMC Medical Research Methodology, 3, 21. doi:10.1186/1471-2288-3-21 .
Campbell, I. (2007). Chi-squared and Fisher-Irwin tests of two-by-two tables with small sample recommendations. Statistics in Medicine, 26(19), 3661–3675. doi:10.1002/sim.2832 .
Katz, D., Baptista, J., Azen, S. P., & Pike, M. C. (1978). Obtaining confidence intervals for the risk ratio in cohort studies. Biometrics, 34(3), 469–474. doi:10.2307/2530610 .
Woolf, B. (1955). On estimating the relation between blood group and disease. Annals of Human Genetics, 19(4), 251–253. doi:10.1111/j.1469-1809.1955.tb01348.x .
Zou, G. (2004). A modified Poisson regression approach to prospective studies with binary data. American Journal of Epidemiology, 159(7), 702–706. doi:10.1093/aje/kwh090 .
See also
tb() for a single cross-tabulation, add_effect() and test()
to modify the table, journal_style() for formatting, and
simtablr_references for all references cited by SimtablR.
Examples
# Describe the whole cohort
table1(epitabl, c(age, sex, smoking, bmi))
#> Characteristic Overall (N=1500)
#> ---------------------------------------------------------------
#> Age at index presentation (years) [Mean (SD)] 62.0 (13.2)
#> Sex recorded for clinical assessment, n (%)
#> Female 686 (45.7%)
#> Male 814 (54.3%)
#> Smoking status, n (%)
#> Never 742 (49.5%)
#> Former 469 (31.3%)
#> Current 289 (19.3%)
#> Body mass index (kg/m2) [Mean (SD)] 27.8 (4.9)
#> Missing 63
# Compare groups, with p-values
table1(epitabl, c(age, sex, smoking), by = adjudicated_acs, flags = "p")
#> `p` now adds the p-value column (SimtablR 3.0); use `perc` or `cell` for total
#> percentages.
#> Characteristic Overall (N=1500) No (N=928) Yes (N=572) P-value
#> --------------------------------------------------------------------------------------------------
#> Age at index presentation (years) [Mean (SD)] 62.0 (13.2) 60.5 (12.8) 64.4 (13.4) <0.001
#> Sex recorded for clinical assessment, n (%) 0.021
#> Female 686 (45.7%) 446 (48.1%) 240 (42.0%)
#> Male 814 (54.3%) 482 (51.9%) 332 (58.0%)
#> Smoking status, n (%) 0.001
#> Never 742 (49.5%) 474 (51.1%) 268 (46.9%)
#> Former 469 (31.3%) 302 (32.5%) 167 (29.2%)
#> Current 289 (19.3%) 152 (16.4%) 137 (24.0%)
#>
#> Tests: Welch t-test; N-1 chi-squared; Pearson's Chi-squared test
#> ℹ Methodological guidance
#> 3 unadjusted p-values are reported in this table. Pipe into test(p.adjust =
#> 'holm') for family-wise control or test(p.adjust = 'BH') for
#> false-discovery-rate control.
#> Run simtablr_guidance("off") separately before printing to hide advice.
# Crude and age- and sex-adjusted prevalence ratios
table1(
epitabl, c(smoking, diabetes, hypertension),
by = adjudicated_acs, measure = "PR", adjust = c(age, sex)
)
#> Characteristic Overall (N=1500) No (N=928) Yes (N=572) PR (95% CI) Adjusted PR (95% CI)
#> --------------------------------------------------------------------------------------------------------------------
#> Smoking status, n (%)
#> Never 742 (49.5%) 474 (51.1%) 268 (46.9%) 1.00 (Ref) 1.00 (Ref)
#> Former 469 (31.3%) 302 (32.5%) 167 (29.2%) 0.99 (0.84 - 1.15) 0.98 (0.84 - 1.14)
#> Current 289 (19.3%) 152 (16.4%) 137 (24.0%) 1.31 (1.12 - 1.53) 1.26 (1.08 - 1.47)
#> History of diabetes, n (%)
#> No 1127 (75.1%) 735 (79.2%) 392 (68.5%) 1.00 (Ref) 1.00 (Ref)
#> Yes 373 (24.9%) 193 (20.8%) 180 (31.5%) 1.39 (1.22 - 1.58) 1.32 (1.16 - 1.50)
#> History of hypertension, n (%)
#> No 732 (48.8%) 463 (49.9%) 269 (47.0%) 1.00 (Ref) 1.00 (Ref)
#> Yes 768 (51.2%) 465 (50.1%) 303 (53.0%) 1.07 (0.94 - 1.22) 0.98 (0.86 - 1.11)
# Add standardized mean differences after the fact
table1(epitabl, c(age, sex), by = adjudicated_acs, test = TRUE) |>
test(smd = TRUE)
#> Characteristic Overall (N=1500) No (N=928) Yes (N=572) SMD P-value
#> --------------------------------------------------------------------------------------------------------
#> Age at index presentation (years) [Mean (SD)] 62.0 (13.2) 60.5 (12.8) 64.4 (13.4) 0.30 <0.001
#> Sex recorded for clinical assessment, n (%) 0.12 0.021
#> Female 686 (45.7%) 446 (48.1%) 240 (42.0%)
#> Male 814 (54.3%) 482 (51.9%) 332 (58.0%)
#>
#> Tests: Welch t-test; N-1 chi-squared
#> ℹ Methodological guidance
#> 2 additional methodological notes hidden. Run simtablr_guidance("teaching")
#> separately before printing to show all advice.