Skip to contents

This example combines a few files bundled with OpenSpecy into a small reference library. Real library builds usually need larger lookup tables and more curation, but the same helper functions apply.

Read And Combine Spectra

c_spec() and build_lib() default to the widest range represented by the source spectra at resolution 6. Values outside each source’s original range are kept as NA, so useful spectral regions are not discarded. This compact vignette example uses the shared overlapping range to keep the rendered output small.

mini_files <- c(
  read_extdata("raman_hdpe.csv"),
  read_extdata("ftir_ldpe_soil.asp"),
  read_extdata("raman_atacamit.spc")
)

mini_sources <- lapply(mini_files, read_any)
mini_sources <- lapply(mini_sources, function(x) {
  x$metadata$intensity_units <- "absorbance"
  attr(x, "intensity_unit") <- "absorbance"
  x
})
mini_raw <- c_spec(mini_sources, range = "common", res = 6)

check_OpenSpecy(mini_raw)
#> [1] TRUE
dim(mini_raw$spectra)
#> [1] 67  3
mini_raw$metadata[, "file_name", with = FALSE]
#>             file_name
#>                <char>
#> 1:     raman_hdpe.csv
#> 2: ftir_ldpe_soil.asp
#> 3: raman_atacamit.spc

Create And Fill A Lookup

Lookup templates help users see which metadata values need curation. When path is not supplied, the template is returned as a data.table; users can write it to CSV by supplying path. Lookup joins use exact values, so edit the template until the key column matches the object metadata.

template <- make_lib_lookup_template(
  mini_raw,
  columns = "file_name",
  add = c("material", "material_type")
)

template
#>             file_name material material_type
#>                <char>   <char>        <char>
#> 1:     raman_hdpe.csv     <NA>          <NA>
#> 2: ftir_ldpe_soil.asp     <NA>          <NA>
#> 3: raman_atacamit.spc     <NA>          <NA>

For this small example, create the lookup table directly in R. The same shape could come from a CSV edited outside R.

lookup <- data.table::data.table(
  file_name = basename(mini_files),
  library_type = c("example", "example", "example"),
  spectrum_type = c("raman", "ftir", "raman"),
  material = c("hdpe", "ldpe in soil", "atacamite")
)

hierarchy <- data.table::data.table(
  material = c("hdpe", "ldpe in soil", "atacamite"),
  material_class = c("polyethylene", "polyethylene", "copper mineral"),
  material_type = c("plastic", "plastic", "mineral")
)

join_lib_metadata(mini_raw, lookup, by = "file_name",
                  require_complete = TRUE)$metadata[
  , c("file_name", "library_type", "spectrum_type", "material"),
  with = FALSE
]
#>             file_name library_type spectrum_type     material
#>                <char>       <char>        <char>       <char>
#> 1:     raman_hdpe.csv      example         raman         hdpe
#> 2: ftir_ldpe_soil.asp      example          ftir ldpe in soil
#> 3: raman_atacamit.spc      example         raman    atacamite

Build A Mini Library

build_lib() can run the ordinary lookup and material hierarchy joins before applying named recipes. Recipe names become names in the returned list. Empty recipes keep the merged spectra unchanged; other recipe lists are passed to process_spec(). Missing values and processing attributes are handled automatically. Metadata column names are also cleaned to lowercase underscore names. Known aliases are coalesced using an editable lookup table. Variants that differ only by underscores or one terminal plural s match automatically.

Before sources are merged, build_lib() also converts declared reflectance and transmittance spectra to absorbance by default. A nonempty attr(x, "intensity_unit") is the primary truth for the whole object; otherwise, metadata$intensity_units is evaluated spectrum by spectrum. Unknown or missing units are left unchanged with a warning. Use convert_intensity = FALSE when unit handling has already been completed outside the builder.

The normal input is one or more file paths; one OpenSpecy or a list of OpenSpecy objects supports sources already loaded in memory. An RDS path may store either one object or a list of objects. Other formats are read with read_any(). Large same-axis source lists are prepared in bulk, which is much faster for legacy RDS files that store many one-spectrum objects. Named stages and elapsed time are reported while the library is built; use progress = FALSE when quiet output is preferred. Supplying restrict_range_args triggers the existing restrict_range() operation before deduplication and recipes; multiple retained ranges can exclude a known silent region without custom workflow code.

name_lookup <- lib_metadata_name_lookup(
  project_code = c("campaign id", "study code"),
  regex = list(instrument_mode = "^method_[0-9]+$")
)
name_lookup[
  canonical_name %in% c("material_color", "number_of_accumulations")
]
#>             canonical_name             source_name  regex
#>                     <char>                  <char> <char>
#> 1:          material_color          material_color   <NA>
#> 2:          material_color                   color   <NA>
#> 3:          material_color                  colour   <NA>
#> 4: number_of_accumulations number_of_accumulations   <NA>
#> 5: number_of_accumulations  number_of_sample_scans   <NA>
#> 6: number_of_accumulations           coadded_scans   <NA>
lib_clean_name(c("User Name", "Laser (%)", "Method...3"))
#> [1] "user_name"  "laser_perc" "method_3"

Named arguments add exact aliases to the defaults, while regex adds patterns evaluated against cleaned names. Overlapping regex patterns produce an error that identifies the source column and matching rules. Pass the result as metadata_name_lookup. Set clean_metadata_values = TRUE in build_lib() or clean_values = TRUE in lib_clean_metadata() to lowercase, trim, and ASCII-normalize character metadata values before joins. Ordinary and hierarchical joins run whenever their corresponding lookup input is non-NULL. spectrum_identity receives one additional deterministic cleanup: recognizable paths are reduced to their basename and trailing file extensions supported by read_any() are removed. Numeric OPUS extensions include .10 and any other terminal period followed only by digits. Exact lookup keys receive the same cleanup. Each built library records changed values and counts in its spectrum_identity_cleanup_report attribute. Keep flexible class patterns in a separate table and apply them afterward with predict_class_reference(). Automatic ordinary lookups use the single shared column with overlapping values and unique lookup keys. Lookups with no usable shared key are skipped with a message; lookups with multiple usable shared keys are treated as ambiguous, so wrap a lookup as list(lookup = table, by = "key") when an explicit key is needed. Named by vectors map metadata names to different lookup names. Canonical source metadata is standardized before these external joins. Every spectrum receives library_name from a populated organization, otherwise from user_name; this is the canonical source-library grouping used by build assessments. A blank organization is still filled from the reviewed internal user_name alias for the existing type lookup. The older lookup-level fallback_by field is deprecated. Set fill_only = TRUE to fill blank values such as library_type or spectrum_type without replacing populated source metadata.

mini_libs <- build_lib(
  mini_files,
  recipes = list(
    raw = list(),
    derivative = list(
      conform_spec = FALSE,
      smooth_intens = TRUE,
      smooth_intens_args = list(window = 15, derivative = 1),
      make_rel = TRUE
    ),
    nobaseline = list(
      conform_spec = FALSE,
      smooth_intens = FALSE,
      subtr_baseline = TRUE,
      make_rel = TRUE
    )
  ),
  metadata_lookups = lookup,
  material_hierarchy = hierarchy,
  clean_metadata_values = TRUE,
  convert_intensity = FALSE,
  assess = TRUE,
  dedupe = FALSE
)
#> build_lib [0.0s]: starting
#> build_lib [0.0s]: reading path 1/3 (raman_hdpe.csv)
#> build_lib [0.2s]: reading path 2/3 (ftir_ldpe_soil.asp)
#> build_lib [0.4s]: reading path 3/3 (raman_atacamit.spc)
#> build_lib [0.7s]: streamed 3 source object(s) from 3 path(s)
#> build_lib [0.7s]: merging 3 source object(s) with c_spec()
#> build_lib [0.9s]: joining metadata lookup 1/1
#> build_lib [0.9s]: metadata lookup 1/1 complete (matched=3; unmatched=0)
#> build_lib [0.9s]: joining the material hierarchy
#> build_lib [0.9s]: material hierarchy complete (unmatched=0)
#> build_lib [0.9s]: processing recipe 1/3 (raw)
#> build_lib [0.9s]: calculating signal-to-noise (raw)
#> build_lib [0.9s]: assessing spectra (raw)
#> build_lib [0.9s]: processing recipe 2/3 (derivative)
#> build_lib [0.9s]: calculating signal-to-noise (derivative)
#> build_lib [0.9s]: assessing spectra (derivative)
#> build_lib [0.9s]: processing recipe 3/3 (nobaseline)
#> build_lib [0.9s]: calculating signal-to-noise (nobaseline)
#> build_lib [0.9s]: assessing spectra (nobaseline)
#> build_lib [0.9s]: complete

names(mini_libs)
#> [1] "raw"        "derivative" "nobaseline"
check_OpenSpecy(mini_libs$raw)
#> [1] TRUE
check_OpenSpecy(mini_libs$derivative)
#> [1] TRUE
attr(mini_libs$derivative, "derivative_order")
#> [1] "1"
attr(mini_libs$nobaseline, "baseline")
#> [1] "nobaseline"
mini_libs$raw$metadata[
  , .(file_name, material, material_class, material_type, sn,
      assessment_flag, assessment_checks)
]
#>             file_name     material material_class material_type        sn
#>                <char>       <char>         <char>        <char>     <num>
#> 1:     raman_hdpe.csv         hdpe   polyethylene       plastic  5.542373
#> 2: ftir_ldpe_soil.asp ldpe in soil   polyethylene       plastic 43.439904
#> 3: raman_atacamit.spc    atacamite copper mineral       mineral  5.220850
#>    assessment_flag assessment_checks
#>             <lgcl>            <char>
#> 1:            TRUE    missing_values
#> 2:            TRUE    missing_values
#> 3:            TRUE    missing_values

prune_lib() performs the spectrum-supported cleanup used before medoid or model creation. Within the same spectral-technique pool, other may be reassigned to the nearest established class, other plastic only to a candidate whose material_type is plastic, and other material only to organic matter or mineral. The matched material_type is copied with the class. It then evaluates material classes from largest to smallest using bounded correlation blocks. FTIR and NIR share a candidate pool; Raman is separate; 2200–2420 cm-1 is excluded by default. After generic reassignment, each spectrum_type and material-class group must contain at least min_n spectra. Smaller groups are removed before correlation pruning, while groups exactly at the threshold are retained whole and larger groups are never reduced below it. Constant or unmatched spectra in eligible classes are retained, and deterministic IDs resolve correlation ties. Use return = "report" for the retained IDs, frozen class schedule, reassignment correlations, spectrum-level removals, and excluded_classes, which reports observed support and the number of additional or reassigned spectra needed. End-to-end builds combine these class rows in assessments$cleanup$summary with the recipe name. build_lib(prune = ...) maps recipe names to prune_lib() argument lists, so each selected spectral representation is pruned independently. Resolved class/type groups below min_n are reassigned as a whole to the most-correlated established class in the same technique pool and material type; they are dropped only when no eligible correlated destination exists. The pruning assessment records the destination, mean class correlation, action, and reason. In the official workflow, an otherwise unresolved identity is temporarily labeled other. The default remove_other = TRUE removes blank identities and the unresolved literal other class before quality control. Reviewed broad other plastic and other material categories stay in the reference libraries and remain eligible for constrained reassignment during derivative/no-baseline pruning. The builder keeps reviewed identifiers, source metadata, prior labels, reasons, actions, and typed before/after counts in assessments$cleanup$summary. The separate assessments$cleanup$dropped_spectrum_identities table contains only the sorted distinct source identities removed anywhere in cleanup. Set remove_other = FALSE to retain generic rows for the constrained semisupervised prune_lib() reassignment described above.

pruned <- prune_lib(
  mini_libs$derivative,
  min_n = 1,
  return = "report",
  progress = FALSE
)
pruned$summary
#>    before after reassigned classes_excluded threshold_removed removed
#>     <int> <int>      <int>            <int>             <int>   <int>
#> 1:      3     3          0                0                 0       0
pruned$schedule
#>        pool material_class initial_n is_protected schedule_order
#>      <char>         <char>     <int>       <lgcl>          <int>
#> 1: ftir_nir   polyethylene         1        FALSE              1
#> 2:    raman copper mineral         1        FALSE              1
#> 3:    raman   polyethylene         1        FALSE              2
pruned$excluded_classes
#> Empty data.table (0 rows and 11 cols): spectrum_type,pool,material_class,observed_n,minimum_spectra,shortfall...

Official Reference-Library Workflow

The version-controlled workflows/OpenSpecy_reference_library.R script retains the source and output path finders and passes those paths to one build_lib() call. The complete visual flow, including checkpoint reuse and old/new assessments, is maintained in .specify/memory/build-lib-diagram.html. Canonically named exact-class, regex-class, library-type, material-hierarchy, and known-bad-ID tables live under workflows/data/. Raw-data corrections are completed externally before this workflow runs. The source-library paths and output directory are explicit inputs. When workflow_data is not supplied, build_lib() looks for the helper tables under data/ beside the calling script, then under data/ or workflows/data/ in the current working directory. Calling build_lib() with no x now stops with an actionable error instead of guessing external paths. The official build derives library_name from organization first and user name second, fills a missing organization from user_name before one source-type lookup, and asserts complete library/spectrum-type coverage. Per-recipe retention rows in assessments$cleanup$summary report counts at preparation, exclusion, deduplication, broad-category review, quality control, pruning, post-transform, and final partition stages. A completely dropped source names the first empty stage and its reason. The exact class lookup runs first; predict_class_reference() then evaluates the separate regex table only for blank materials. Exact/regex overlaps are audited without overwriting exact values, and conflicting regex predictions stop for review. Material classes use concise polymer-family names, with polyethylene and polypropylene separated and polyhydroxy(meth)acrylates retaining its chemically meaningful optional group notation. Remaining uncertain identities follow the selected generic-row review policy. Derivative and nobaseline spectra pass three quality gates before pruning: FTIR CO2 is flattened when the 2200–2420 maximum divided by the 2420–2550 silent-region maximum is greater than two; high-tail detection ignores NA padding, trims each spectrum on its finite support, and drops failed corrections; and finite running SNR below two (or unavailable SNR) is removed. Raw remains unpruned. Each completed library, medoid, model, and assessment component is written with a SHA-256 manifest under output_dir/checkpoints. reuse = TRUE loads it only when the source/lookup signatures, arguments, package version, and builder contract still match. Validated output is promoted into a versioned release directory, which contains the seven legacy RDS names, explicit algorithm-named model RDS files, one aggregate build RDS, and a release manifest. Existing versioned payloads are immutable: a different hash stops promotion.

The returned object has four top-level elements: libraries, medoids, models, and assessments. Libraries are recipe-first and then keyed by ftir, raman, or nir: full Raman uses 200–4000, FTIR uses 400–4000, and NIR uses 4000–12000 cm-1. FTIR/Raman medoids and models use 800–3200; the NIR identification interval is derived from the longest region with at least 90% finite coverage (falling back to available coverage) and recorded on the object. All-blank metadata columns are dropped independently by type. Each spectrum must observe at least 10% of its type-specific identification axis to enter medoid selection or model fitting. PAM operates on a temporary relative matrix where each spectrum’s missing values are filled by that spectrum’s finite mean. Selected medoid IDs are then pulled from the original range-restricted object, so published medoids retain their original NA positions. Groups with more than 3,000 spectra use five deterministic 1,000-spectrum PAM samples; each candidate set is scored against the complete group with the same correlation distance, avoiding a full oversized dissimilarity matrix. models is algorithm-first (logistic_regression or random_forest), then recipe and type. Logistic regression retains derivative/nobaseline medoid training and fills restored gaps with the finite mean at each wavenumber across training medoids. Random forests use the complete eligible raw, derivative, and nobaseline type libraries, relative normalization, wavenumber-mean filling, inverse-frequency balanced case sampling, 500 probability trees, permutation importance, and all available ranger threads. Balanced sampling improves minority representation in each bootstrap without also applying a class-vote correction. Because this resampling can make out-of-bag accuracy optimistic, treat OOB results as fit diagnostics. The release assessment separately applies each deployed model to its complete corresponding source dataset. train_spec_model() exposes both engines; build_model_lib() remains a compatible wrapper. Each model stores its filler, support audit, training diagnostics, and one tidy tests table. Logistic lambda selection remains maximum out-of-fold macro class accuracy with its calibrated alpha, no-intercept, grouped multinomial, and weighting settings unchanged.

Full assessment uses every candidate and downloaded legacy artifact. Candidate and legacy artifacts independently use a seeded approximately ten-percent class/type-stratified sample. Rows sharing a physical identifier or exact spectral content are one group and cannot cross training and test. This source-local policy tolerates taxonomy evolution without fuzzy cross-version class matching; its results describe each artifact on its own source population and are not paired causal deltas. Held-out groups are removed from full reference libraries before their identification assessment. Medoid artifacts are instead tested as deployed: each medoid library identifies every spectrum in its complete corresponding processed library (for example, medoid_derivative.rds searches derivative.rds). Existing logistic and random-forest artifacts likewise identify their complete corresponding source libraries once; assessment does not select new medoids or call train_spec_model(). match_spec() uses stored fillers for partial spectra. Macro class accuracy is the primary metric, followed by evaluated coverage, overall accuracy, and clear old/new assess_spec(report = "all") shifts. The review surface contains at most ten tables nested under cleanup, ref_lib, medoid, model, and functionality. Old and new values occupy adjacent columns; accuracy, misidentification counts, absolute diagnostic correlations, and warning/error rate shifts are sorted from largest to smallest. Pass rows are omitted from quality shifts. Machine-level split and row-test evidence is available through attr(reference_library_build$assessments, "evidence"). Published accuracy tables contain aggregate overall and macro metrics, not per-class accuracy rows. A versioned release writes all global cleanup, quality, pruning, model, and comparison evidence to assessments.rds; its library, medoid, and model files contain only runtime data and prediction state. The accompanying reference_library_build.rds is a lightweight index that can be supplied to rebuild_lib_artifacts(). The smaller 1,000-spectrum run in the manual benchmark is development-only and is not part of build_lib() or final acceptance evidence.

The standard call therefore has this shape:

reference_library_build <- build_lib(
  files,
  output_dir = output_dir,
  previous_library_dir = "system",
  remove_other = TRUE
)

After upstream libraries have already completed, rebuild only the downstream artifacts into a new explicit output location:

updated_reference_library_build <- rebuild_lib_artifacts(
  completed_build_dir,
  output_dir = downstream_output_dir,
  previous_library_dir = "system"
)

completed_build_dir may instead be a reference_library_build.rds, a libraries.rds checkpoint, or the corresponding in-memory build object.

The script is excluded from package builds but remains available in the GitHub repository so library releases can be reviewed and reproduced.