mice
mice can finally predict, not just estimate, from multiply imputed data.
A side-by-side editorial comparison of bulkreadr and sdcMicro — release velocity, themes, recent moves, and the top alternatives to consider.
A bulk file reader became a labelled-survey-data toolkit, then went quiet
bulkreadr started as a way to read many files at once and turned into tooling for labelled survey data: SPSS and Stata importers that convert labelled variables to factors, generate_dictionary() for building data dictionaries, look_for() for searching variable descriptions, and imputation helpers. The most recent release does the opposite of adding — it pulls inspect_na() in-house to drop an external dependency.
A 20-year anonymization toolbox now has a language model inside its refinement loop.
sdcMicro is the reference R implementation of statistical disclosure control — k-anonymity, local suppression, PRAM, microaggregation, record swapping — used by national statistical offices, with a Shiny GUI (sdcApp) as its second face. The feed shows a long GUI-maintenance era through 2018-2022 and then a gap, and the package that reappears in 5.8.2 has an AI_applyAnonymization() workflow and a query_llm() helper that the older entries know nothing about. The July release tunes that loop rather than introducing it.
bulkreadr started as a way to read many files at once and turned into tooling for labelled survey data: SPSS and Stata importers that convert labelled variables to factors, generate_dictionary() for building data dictionaries, look_for() for searching variable descriptions, and imputation helpers. The most recent release does the opposite of adding — it pulls inspect_na() in-house to drop an external dependency.
Growth came in a burst across 2023, slowed to one release a year, and has now turned inward. The 2023 cadence added a format or a labelled-data function every few weeks; 2025 added a single Excel-to-CSV exporter; 2026 removed a dependency. The GitHub notes are cumulative — each release restates every prior version's changelog — which makes the feed look busier than the work is.
With inspectdf gone, the remaining Suggests-level dependencies are the obvious next targets for the same treatment. Nothing in these entries points to a new file format or a return to the 2023 pace.
sdcMicro is the reference R implementation of statistical disclosure control — k-anonymity, local suppression, PRAM, microaggregation, record swapping — used by national statistical offices, with a Shiny GUI (sdcApp) as its second face. The feed shows a long GUI-maintenance era through 2018-2022 and then a gap, and the package that reappears in 5.8.2 has an AI_applyAnonymization() workflow and a query_llm() helper that the older entries know nothing about. The July release tunes that loop rather than introducing it.
Two threads run in parallel. The visible one is the LLM-assisted anonymization path maturing: 5.8.2 gives its refinement loop early stopping via tol and patience so it stops when the combined utility score plateaus instead of burning all max_iter rounds, and teaches query_llm() to drop the temperature parameter for reasoning models that reject it. The other is unglamorous statistical correctness — a distinct l-diversity computation fixed for NAs in key variables, with the C++ simplified and tests added. The release also ships reproducibility scripts for a SoftwareX paper, which suggests the AI path is being written up rather than quietly trialled.
The provider-compatibility fix is reactive — a parameter dropped because one model family rejected it — so expect more of the same as query_llm() meets other backends. Given tol and patience were added to stop wasted iterations, cost or runtime of the refinement loop is the live concern, and further controls on it are the likeliest next move.
Other Infra & APIs products tracked by Sparkpulse, ranked by recent ship velocity. Each card links to a full editorial trajectory and lets you pivot into a head-to-head comparison with either bulkreadr or sdcMicro.
mice can finally predict, not just estimate, from multiply imputed data.
A market-microstructure toolkit that keeps adding estimators as the papers land.
A vowel-analysis package trimming dependencies after an email address got it archived.
The R half of the EMU speech database system, fixing what was quietly broken.
A Bayesian model-averaging package spending its 2.0 on memory, not methods.
tidyplots keeps rebuilding its own foundations rather than layering around them.
See all bulkreadr alternatives → · See all sdcMicro alternatives →
Latest ship moves from both products, interleaved chronologically. ⚡ = editorial spark.
They serve adjacent needs but don't currently overlap on shipped themes. bulkreadr and sdcMicro are shipping at a similar cadence (velocity 0.0 vs 0.0, both within Sparkpulse's "active" band). See the at-a-glance table above for a side-by-side breakdown of velocity, recent sparks, and editorial themes.
Sparkpulse doesn't pick a winner — we score release velocity, not feature parity. bulkreadr and sdcMicro are shipping at a similar cadence (velocity 0.0 vs 0.0, both within Sparkpulse's "active" band). For your specific use case, the alternatives sections above list other Infra & APIs products to evaluate alongside.
Top bulkreadr alternatives in Infra & APIs are ranked by recent ship velocity. Browse the "bulkreadr alternatives" section above for the current picks, or visit /alternatives/bulkreadr for the full list with editorial commentary on each.
Top sdcMicro alternatives in Infra & APIs are ranked by recent ship velocity. Browse the "sdcMicro alternatives" section above for the current picks, or visit /alternatives/sdcmicro for the full list with editorial commentary on each.