Build a Mini Reference Library
Win Cowger
2026-09-24
Source:vignettes/library-builder.Rmd
library-builder.RmdThis example combines a few files bundled with OpenSpecy into a small reference library. Real library builds usually need larger lookup tables and more curation, but the same helper functions apply.
Read And Combine Spectra
c_spec() and build_lib() default to the
widest range represented by the source spectra at resolution 6. Values
outside each source’s original range are kept as NA, so
useful spectral regions are not discarded. This compact vignette example
uses the shared overlapping range to keep the rendered output small.
mini_files <- c(
read_extdata("raman_hdpe.csv"),
read_extdata("ftir_ldpe_soil.asp"),
read_extdata("raman_atacamit.spc")
)
mini_sources <- lapply(mini_files, read_any)
mini_sources <- lapply(mini_sources, function(x) {
x$metadata$intensity_units <- "absorbance"
attr(x, "intensity_unit") <- "absorbance"
x
})
mini_raw <- c_spec(mini_sources, range = "common", res = 6)
check_OpenSpecy(mini_raw)
#> [1] TRUE
dim(mini_raw$spectra)
#> [1] 67 3
mini_raw$metadata[, "file_name", with = FALSE]
#> file_name
#> <char>
#> 1: raman_hdpe.csv
#> 2: ftir_ldpe_soil.asp
#> 3: raman_atacamit.spcCreate And Fill A Lookup
Lookup templates help users see which metadata values need curation.
When path is not supplied, the template is returned as a
data.table; users can write it to CSV by supplying
path. Lookup joins use exact values, so edit the template
until the key column matches the object metadata.
template <- make_lib_lookup_template(
mini_raw,
columns = "file_name",
add = c("material", "material_type")
)
template
#> file_name material material_type
#> <char> <char> <char>
#> 1: raman_hdpe.csv <NA> <NA>
#> 2: ftir_ldpe_soil.asp <NA> <NA>
#> 3: raman_atacamit.spc <NA> <NA>For this small example, create the lookup table directly in R. The same shape could come from a CSV edited outside R.
lookup <- data.table::data.table(
file_name = basename(mini_files),
library_type = c("example", "example", "example"),
spectrum_type = c("raman", "ftir", "raman"),
material = c("hdpe", "ldpe in soil", "atacamite")
)
hierarchy <- data.table::data.table(
material = c("hdpe", "ldpe in soil", "atacamite"),
material_class = c("polyethylene", "polyethylene", "copper mineral"),
material_type = c("plastic", "plastic", "mineral")
)
join_lib_metadata(mini_raw, lookup, by = "file_name",
require_complete = TRUE)$metadata[
, c("file_name", "library_type", "spectrum_type", "material"),
with = FALSE
]
#> file_name library_type spectrum_type material
#> <char> <char> <char> <char>
#> 1: raman_hdpe.csv example raman hdpe
#> 2: ftir_ldpe_soil.asp example ftir ldpe in soil
#> 3: raman_atacamit.spc example raman atacamiteBuild A Mini Library
build_lib() can run the ordinary lookup and material
hierarchy joins before applying named recipes. Recipe names become names
in the returned list. Empty recipes keep the merged spectra unchanged;
other recipe lists are passed to process_spec(). Missing
values and processing attributes are handled automatically. Metadata
column names are also cleaned to lowercase underscore names. Known
aliases are coalesced using an editable lookup table. Variants that
differ only by underscores or one terminal plural s match
automatically.
Before sources are merged, build_lib() also converts
declared reflectance and transmittance spectra to absorbance by default.
A nonempty attr(x, "intensity_unit") is the primary truth
for the whole object; otherwise, metadata$intensity_units
is evaluated spectrum by spectrum. Unknown or missing units are left
unchanged with a warning. Use convert_intensity = FALSE
when unit handling has already been completed outside the builder.
The normal input is one or more file paths; one
OpenSpecy or a list of OpenSpecy objects
supports sources already loaded in memory. An RDS path may store either
one object or a list of objects. Other formats are read with
read_any(). Large same-axis source lists are prepared in
bulk, which is much faster for legacy RDS files that store many
one-spectrum objects. Named stages and elapsed time are reported while
the library is built; use progress = FALSE when quiet
output is preferred. Supplying restrict_range_args triggers
the existing restrict_range() operation before
deduplication and recipes; multiple retained ranges can exclude a known
silent region without custom workflow code.
name_lookup <- lib_metadata_name_lookup(
project_code = c("campaign id", "study code"),
regex = list(instrument_mode = "^method_[0-9]+$")
)
name_lookup[
canonical_name %in% c("material_color", "number_of_accumulations")
]
#> canonical_name source_name regex
#> <char> <char> <char>
#> 1: material_color material_color <NA>
#> 2: material_color color <NA>
#> 3: material_color colour <NA>
#> 4: number_of_accumulations number_of_accumulations <NA>
#> 5: number_of_accumulations number_of_sample_scans <NA>
#> 6: number_of_accumulations coadded_scans <NA>
lib_clean_name(c("User Name", "Laser (%)", "Method...3"))
#> [1] "user_name" "laser_perc" "method_3"Named arguments add exact aliases to the defaults, while
regex adds patterns evaluated against cleaned names.
Overlapping regex patterns produce an error that identifies the source
column and matching rules. Pass the result as
metadata_name_lookup. Set
clean_metadata_values = TRUE in build_lib() or
clean_values = TRUE in lib_clean_metadata() to
lowercase, trim, and ASCII-normalize character metadata values before
joins. Ordinary and hierarchical joins run whenever their corresponding
lookup input is non-NULL. spectrum_identity
receives one additional deterministic cleanup: recognizable paths are
reduced to their basename and trailing file extensions supported by
read_any() are removed. Numeric OPUS extensions include
.10 and any other terminal period followed only by digits.
Exact lookup keys receive the same cleanup. Each built library records
changed values and counts in its
spectrum_identity_cleanup_report attribute. Keep flexible
class patterns in a separate table and apply them afterward with
predict_class_reference(). Automatic ordinary lookups use
the single shared column with overlapping values and unique lookup keys.
Lookups with no usable shared key are skipped with a message; lookups
with multiple usable shared keys are treated as ambiguous, so wrap a
lookup as list(lookup = table, by = "key") when an explicit
key is needed. Named by vectors map metadata names to
different lookup names. Canonical source metadata is standardized before
these external joins. Every spectrum receives library_name
from a populated organization, otherwise from
user_name; this is the canonical source-library grouping
used by build assessments. A blank organization is still
filled from the reviewed internal user_name alias for the
existing type lookup. The older lookup-level fallback_by
field is deprecated. Set fill_only = TRUE to fill blank
values such as library_type or spectrum_type
without replacing populated source metadata.
mini_libs <- build_lib(
mini_files,
recipes = list(
raw = list(),
derivative = list(
conform_spec = FALSE,
smooth_intens = TRUE,
smooth_intens_args = list(window = 15, derivative = 1),
make_rel = TRUE
),
nobaseline = list(
conform_spec = FALSE,
smooth_intens = FALSE,
subtr_baseline = TRUE,
make_rel = TRUE
)
),
metadata_lookups = lookup,
material_hierarchy = hierarchy,
clean_metadata_values = TRUE,
convert_intensity = FALSE,
assess = TRUE,
dedupe = FALSE
)
#> build_lib [0.0s]: starting
#> build_lib [0.0s]: reading path 1/3 (raman_hdpe.csv)
#> build_lib [0.2s]: reading path 2/3 (ftir_ldpe_soil.asp)
#> build_lib [0.4s]: reading path 3/3 (raman_atacamit.spc)
#> build_lib [0.7s]: streamed 3 source object(s) from 3 path(s)
#> build_lib [0.7s]: merging 3 source object(s) with c_spec()
#> build_lib [0.9s]: joining metadata lookup 1/1
#> build_lib [0.9s]: metadata lookup 1/1 complete (matched=3; unmatched=0)
#> build_lib [0.9s]: joining the material hierarchy
#> build_lib [0.9s]: material hierarchy complete (unmatched=0)
#> build_lib [0.9s]: processing recipe 1/3 (raw)
#> build_lib [0.9s]: calculating signal-to-noise (raw)
#> build_lib [0.9s]: assessing spectra (raw)
#> build_lib [0.9s]: processing recipe 2/3 (derivative)
#> build_lib [0.9s]: calculating signal-to-noise (derivative)
#> build_lib [0.9s]: assessing spectra (derivative)
#> build_lib [0.9s]: processing recipe 3/3 (nobaseline)
#> build_lib [0.9s]: calculating signal-to-noise (nobaseline)
#> build_lib [0.9s]: assessing spectra (nobaseline)
#> build_lib [0.9s]: complete
names(mini_libs)
#> [1] "raw" "derivative" "nobaseline"
check_OpenSpecy(mini_libs$raw)
#> [1] TRUE
check_OpenSpecy(mini_libs$derivative)
#> [1] TRUE
attr(mini_libs$derivative, "derivative_order")
#> [1] "1"
attr(mini_libs$nobaseline, "baseline")
#> [1] "nobaseline"
mini_libs$raw$metadata[
, .(file_name, material, material_class, material_type, sn,
assessment_flag, assessment_checks)
]
#> file_name material material_class material_type sn
#> <char> <char> <char> <char> <num>
#> 1: raman_hdpe.csv hdpe polyethylene plastic 5.542373
#> 2: ftir_ldpe_soil.asp ldpe in soil polyethylene plastic 43.439904
#> 3: raman_atacamit.spc atacamite copper mineral mineral 5.220850
#> assessment_flag assessment_checks
#> <lgcl> <char>
#> 1: TRUE missing_values
#> 2: TRUE missing_values
#> 3: TRUE missing_valuesprune_lib() performs the spectrum-supported cleanup used
before medoid or model creation. Within the same spectral-technique
pool, other may be reassigned to the nearest established
class, other plastic only to a candidate whose
material_type is plastic, and
other material only to organic matter or
mineral. The matched material_type is copied
with the class. It then evaluates material classes from largest to
smallest using bounded correlation blocks. FTIR and NIR share a
candidate pool; Raman is separate; 2200–2420 cm-1 is excluded
by default. After generic reassignment, each spectrum_type
and material-class group must contain at least min_n
spectra. Smaller groups are removed before correlation pruning, while
groups exactly at the threshold are retained whole and larger groups are
never reduced below it. Constant or unmatched spectra in eligible
classes are retained, and deterministic IDs resolve correlation ties.
Use return = "report" for the retained IDs, frozen class
schedule, reassignment correlations, spectrum-level removals, and
excluded_classes, which reports observed support and the
number of additional or reassigned spectra needed. End-to-end builds
combine these class rows in assessments$cleanup$summary
with the recipe name. build_lib(prune = ...) maps recipe
names to prune_lib() argument lists, so each selected
spectral representation is pruned independently. Resolved class/type
groups below min_n are reassigned as a whole to the
most-correlated established class in the same technique pool and
material type; they are dropped only when no eligible correlated
destination exists. The pruning assessment records the destination, mean
class correlation, action, and reason. In the official workflow, an
otherwise unresolved identity is temporarily labeled other.
The default remove_other = TRUE removes blank identities
and the unresolved literal other class before quality
control. Reviewed broad other plastic and
other material categories stay in the reference libraries
and remain eligible for constrained reassignment during
derivative/no-baseline pruning. The builder keeps reviewed identifiers,
source metadata, prior labels, reasons, actions, and typed before/after
counts in assessments$cleanup$summary. The separate
assessments$cleanup$dropped_spectrum_identities table
contains only the sorted distinct source identities removed anywhere in
cleanup. Set remove_other = FALSE to retain generic rows
for the constrained semisupervised prune_lib() reassignment
described above.
pruned <- prune_lib(
mini_libs$derivative,
min_n = 1,
return = "report",
progress = FALSE
)
pruned$summary
#> before after reassigned classes_excluded threshold_removed removed
#> <int> <int> <int> <int> <int> <int>
#> 1: 3 3 0 0 0 0
pruned$schedule
#> pool material_class initial_n is_protected schedule_order
#> <char> <char> <int> <lgcl> <int>
#> 1: ftir_nir polyethylene 1 FALSE 1
#> 2: raman copper mineral 1 FALSE 1
#> 3: raman polyethylene 1 FALSE 2
pruned$excluded_classes
#> Empty data.table (0 rows and 11 cols): spectrum_type,pool,material_class,observed_n,minimum_spectra,shortfall...Official Reference-Library Workflow
The version-controlled workflows/OpenSpecy_reference_library.R
script retains the source and output path finders and passes those paths
to one build_lib() call. The complete visual flow,
including checkpoint reuse and old/new assessments, is maintained in .specify/memory/build-lib-diagram.html.
Canonically named exact-class, regex-class, library-type,
material-hierarchy, and known-bad-ID tables live under
workflows/data/. Raw-data corrections are completed
externally before this workflow runs. The source-library paths and
output directory are explicit inputs. When workflow_data is
not supplied, build_lib() looks for the helper tables under
data/ beside the calling script, then under
data/ or workflows/data/ in the current
working directory. Calling build_lib() with no
x now stops with an actionable error instead of guessing
external paths. The official build derives library_name
from organization first and user name second, fills a missing
organization from user_name before one source-type lookup,
and asserts complete library/spectrum-type coverage. Per-recipe
retention rows in assessments$cleanup$summary report counts
at preparation, exclusion, deduplication, broad-category review, quality
control, pruning, post-transform, and final partition stages. A
completely dropped source names the first empty stage and its reason.
The exact class lookup runs first;
predict_class_reference() then evaluates the separate regex
table only for blank materials. Exact/regex overlaps are audited without
overwriting exact values, and conflicting regex predictions stop for
review. Material classes use concise polymer-family names, with
polyethylene and polypropylene separated and
polyhydroxy(meth)acrylates retaining its chemically
meaningful optional group notation. Remaining uncertain identities
follow the selected generic-row review policy. Derivative and nobaseline
spectra pass three quality gates before pruning: FTIR CO2 is flattened
when the 2200–2420 maximum divided by the 2420–2550 silent-region
maximum is greater than two; high-tail detection ignores NA padding,
trims each spectrum on its finite support, and drops failed corrections;
and finite running SNR below two (or unavailable SNR) is removed. Raw
remains unpruned. Each completed library, medoid, model, and assessment
component is written with a SHA-256 manifest under
output_dir/checkpoints. reuse = TRUE loads it
only when the source/lookup signatures, arguments, package version, and
builder contract still match. Validated output is promoted into a
versioned release directory, which contains the seven legacy RDS names,
explicit algorithm-named model RDS files, one aggregate build RDS, and a
release manifest. Existing versioned payloads are immutable: a different
hash stops promotion.
The returned object has four top-level elements:
libraries, medoids, models, and
assessments. Libraries are recipe-first and then keyed by
ftir, raman, or nir: full Raman
uses 200–4000, FTIR uses 400–4000, and NIR uses 4000–12000
cm-1. FTIR/Raman medoids and models use 800–3200; the NIR
identification interval is derived from the longest region with at least
90% finite coverage (falling back to available coverage) and recorded on
the object. All-blank metadata columns are dropped independently by
type. Each spectrum must observe at least 10% of its type-specific
identification axis to enter medoid selection or model fitting. PAM
operates on a temporary relative matrix where each spectrum’s missing
values are filled by that spectrum’s finite mean. Selected medoid IDs
are then pulled from the original range-restricted object, so published
medoids retain their original NA positions. Groups with
more than 3,000 spectra use five deterministic 1,000-spectrum PAM
samples; each candidate set is scored against the complete group with
the same correlation distance, avoiding a full oversized dissimilarity
matrix. models is algorithm-first
(logistic_regression or random_forest), then
recipe and type. Logistic regression retains derivative/nobaseline
medoid training and fills restored gaps with the finite mean at each
wavenumber across training medoids. Random forests use the complete
eligible raw, derivative, and nobaseline type libraries, relative
normalization, wavenumber-mean filling, inverse-frequency balanced case
sampling, 500 probability trees, permutation importance, and all
available ranger threads. Balanced sampling improves minority
representation in each bootstrap without also applying a class-vote
correction. Because this resampling can make out-of-bag accuracy
optimistic, treat OOB results as fit diagnostics. The release assessment
separately applies each deployed model to its complete corresponding
source dataset. train_spec_model() exposes both engines;
build_model_lib() remains a compatible wrapper. Each model
stores its filler, support audit, training diagnostics, and one tidy
tests table. Logistic lambda selection remains maximum
out-of-fold macro class accuracy with its calibrated alpha,
no-intercept, grouped multinomial, and weighting settings unchanged.
Full assessment uses every candidate and downloaded legacy artifact.
Candidate and legacy artifacts independently use a seeded approximately
ten-percent class/type-stratified sample. Rows sharing a physical
identifier or exact spectral content are one group and cannot cross
training and test. This source-local policy tolerates taxonomy evolution
without fuzzy cross-version class matching; its results describe each
artifact on its own source population and are not paired causal deltas.
Held-out groups are removed from full reference libraries before their
identification assessment. Medoid artifacts are instead tested as
deployed: each medoid library identifies every spectrum in its complete
corresponding processed library (for example,
medoid_derivative.rds searches
derivative.rds). Existing logistic and random-forest
artifacts likewise identify their complete corresponding source
libraries once; assessment does not select new medoids or call
train_spec_model(). match_spec() uses stored
fillers for partial spectra. Macro class accuracy is the primary metric,
followed by evaluated coverage, overall accuracy, and clear old/new
assess_spec(report = "all") shifts. The review surface
contains at most ten tables nested under cleanup,
ref_lib, medoid, model, and
functionality. Old and new values occupy adjacent columns;
accuracy, misidentification counts, absolute diagnostic correlations,
and warning/error rate shifts are sorted from largest to smallest. Pass
rows are omitted from quality shifts. Machine-level split and row-test
evidence is available through
attr(reference_library_build$assessments, "evidence").
Published accuracy tables contain aggregate overall and macro metrics,
not per-class accuracy rows. A versioned release writes all global
cleanup, quality, pruning, model, and comparison evidence to
assessments.rds; its library, medoid, and model files
contain only runtime data and prediction state. The accompanying
reference_library_build.rds is a lightweight index that can
be supplied to rebuild_lib_artifacts(). The smaller
1,000-spectrum run in the manual benchmark is development-only and is
not part of build_lib() or final acceptance evidence.
The standard call therefore has this shape:
reference_library_build <- build_lib(
files,
output_dir = output_dir,
previous_library_dir = "system",
remove_other = TRUE
)After upstream libraries have already completed, rebuild only the downstream artifacts into a new explicit output location:
updated_reference_library_build <- rebuild_lib_artifacts(
completed_build_dir,
output_dir = downstream_output_dir,
previous_library_dir = "system"
)completed_build_dir may instead be a
reference_library_build.rds, a libraries.rds
checkpoint, or the corresponding in-memory build object.
The script is excluded from package builds but remains available in the GitHub repository so library releases can be reviewed and reproduced.