Most of our work has resulted in scholarly publications. On this page you can review our publications to get an idea about our work.
Overview of the documentation efforts Gold nanoparticles (GNPs) exhibit unique optical properties governed by localized surface plasmon resonance, enabling applications in biomedicine, sensing, catalysis, and photonics. Accurate … Gold nanoparticles (GNPs) exhibit unique optical properties governed by localized surface plasmon resonance, enabling applications in biomedicine, sensing, catalysis, and photonics. Accurate concentration tracking during fabrication is essential, since particle density strongly affects colloid stability and functional performance. Pulsed laser ablation in liquids (PLAL) has gained recognition as a sustainable route for nanoparticle synthesis, ligand-free, high-purity colloids without chemical by-products. However, monitoring GNP concentration during PLAL typically relies on UV–VIS spectrophotometers, that can be costly, bulky, and difficult to integrate into fabrication workflows. In this work, we present the design, construction, and validation of a compact, open-source alternative for real-time nanoparticle tracking. The device combines a 405 nm laser, a photodiode, and an ESP32-S3 microcontroller for data-acquisition and processing. All design files, firmware, and 3D-printed casing are openly released. Validation experiments show strong agreement with conventional UV–VIS absorption, demonstrating that the proposed device can replace commercial spectrophotometers. By providing a concentration sensor that can be easily integrated in a production chain, this approach reduces the physical footprint and cost of GNPs fabrication facilities, while maintaining reliable monitoring. With a cost below 40 € and modular architecture, the system enables reproducible, accessible, and sustainable monitoring for PLAL and other nanomaterial synthesis processes. В статье рассматриваются особенности применения технологий искусственного интеллекта в англо-узбекском художественном переводе в сопоставлении с переводческой деятельностью профессионального переводчи… В статье рассматриваются особенности применения технологий искусственного интеллекта в англо-узбекском художественном переводе в сопоставлении с переводческой деятельностью профессионального переводчика. Исследование выполнено в русле сравнительного языкознания и современного лингвистического переводоведения с опорой на транскультурный подход, позволяющий рассматривать художественный перевод как процесс переноса не только языкового содержания, но и эстетической, стилистической, культурной и прагматической информации. Материалом исследования послужили 30 фрагментов из произведений Pride and Prejudice, The Great Gatsby, Harry Potter, Animal Farm и The Little Prince, переведённых с использованием ChatGPT, DeepL и Google Translate и сопоставленных с переводами профессионального переводчика. В качестве основных критериев анализа определены семантическая точность, стилистическая совместимость, лингвокультурная эквивалентность и прагматическая адекватность. Отдельно исследованы способы передачи метафор, фразеологических единиц, культурных реалий и стилистических средств. Результаты показывают, что системы искусственного интеллекта способны обеспечивать достаточно высокий уровень семантической точности и лексико-грамматической последовательности, однако при передаче образности, культурной специфики, фразеологических значений и прагматико-эстетической функции художественного текста выявляются определённые переводческие расхождения. Полученные результаты позволяют обосновать необходимость сочетания автоматизированного перевода, сопоставительного анализа и профессиональной постредакции в англо-узбекском художественном переводе. Analysis code for the manuscript "A Multi-Omics Brain-Heart Interactome Maps Signaling Networks Associated with Post-Stroke Cardiac Dysfunction" (prepared for submission to Frontiers in Cell and Devel… Analysis code for the manuscript "A Multi-Omics Brain-Heart Interactome Maps Signaling Networks Associated with Post-Stroke Cardiac Dysfunction" (prepared for submission to Frontiers in Cell and Developmental Biology).
This record contains the complete, reproducible Python analysis pipeline accompanying the manuscript. It covers the matched multi-omics analyses of a mouse MCAO stroke model (brain scRNA-seq, heart scRNA-seq and plasma DIA proteomics from the same animals): brain and heart single-cell RNA-seq processing and marker-based annotation (steps 01-02); per-cell-type expression statistics reported descriptively, consistent with the pooled-library design (step 03); plasma DIA proteomics differential analysis (Welch's t-test on log2 intensities with Benjamini-Hochberg FDR; primary significance rule |log2FC| > 0.5 and q < 0.05) (step 04); the direction-aware, FDR-controlled BOPPS (Brain-to-Organ Plasma Proteomics Score) prioritization framework, including the weight/cutoff sensitivity grid and high-confidence axis calling (step 05); key statistics and supplementary-table generation (steps 06-07); figure panel generation and the final figure assembly (steps 08-10, 15-16); microglial substate analysis (step 11); robustness analyses including leave-one-cell-type-out stability and a label-permutation null model (step 12); two-tier SAA amplification (step 13) and axis-level direction consistency (step 14); and the documented external-validation workflow for the public MCAO datasets GSE174574 and GSE303321 (step 17).
Contents
- 01-17 scripts: numbered analysis scripts (07 is split into four parts, 07_tables_a-d); script-to-output mapping and documented run order in README.md.
- README.md: overview, run order, data layout and the statistical conventions of the study.
- requirements.txt: Python 3.11 environment (scanpy, scrublet, pandas, numpy, scipy, statsmodels, openpyxl, matplotlib, plotly, pymupdf, pillow).
- data/: expected input layout (README.txt); curated scaffold workbooks enabling steps 07-10 to be re-run without reprocessing raw single-cell data; the Figure 5 source composite; and the derived external-validation sheets for Table S46.
Data and reproducibility
Raw single-cell RNA sequencing data have been deposited in the Genome Sequence Archive (GSA) of the National Genomics Data Center under BioProject PRJCA075314. Raw plasma proteomics data have been deposited in the ProteomeXchange Consortium via the iProX partner repository (dataset identifier PXD084849; iProX project ID IPX0019935000). Processed data supporting the findings are provided as Supplementary Tables S1-S46 of the manuscript; public datasets re-analysed for validation are available under GEO accessions GSE174574 and GSE303321. Consistent with the manuscript: single-cell libraries were generated from pooled cell suspensions (one library per group per tissue); single-cell results are descriptive (effect sizes and detection frequencies; no cell-level P-values), and formal statistical inference is restricted to the individually replicated assays; BOPPS is a hypothesis-generating prioritization framework and does not establish directional signaling or causality. The archived Analysis_Code.zip is the same file as the supplementary material submitted with the manuscript (md5 fd53992d41d639ef5de20e432404f81c). Hybrid plasmonic substrates based on graphene oxide (GO) sheets decorated with colloidal gold nanoparticles (NPs) have been tested for surface-enhanced Raman scattering (SERS). By leveraging the … Hybrid plasmonic substrates based on graphene oxide (GO) sheets decorated with colloidal gold nanoparticles (NPs) have been tested for surface-enhanced Raman scattering (SERS). By leveraging the high chemical reactivity of GO and the plasmonic characteristics of gold NPs, these hybrid substrates exhibit improved SERS performance. The morphology and electro-optical properties of the GO-Au NP nanohybrids are thoroughly characterized using scanning probe microscopies, including Kelvin probe force microscopy. Maps of the substrates topography and surface potential are then correlated with the local SERS enhancement factors through co-localized confocal Raman spectroscopy, providing valuable insights into the structure–function relationships of the hybrid plasmonic substrates and contributing to the understanding of SERS mechanisms. The inexpensive and straightforward fabrication process offers a cost-effective and user-friendly approach, while the incorporation of GO sheets provides a scalable and versatile platform, further enhancing the applicability and utility of these substrates. Replication package: simulator, per-topology records and analysis code for learned route scoring in wireless sensor networks. Manuscript submitted for publication.
This archive contains the simulator … Replication package: simulator, per-topology records and analysis code for learned route scoring in wireless sensor networks. Manuscript submitted for publication.
This archive contains the simulator that produced every result in the study, the per-topology records from which every results table and figure is computed, the analysis code, and the protocol card and complete log of the evaluation-only pass on the held-out seeds. The frozen configuration is fingerprinted as 7cf77fea1129bedd. Version 1.1 adds the index-parameterized comparator (configuration hash 760b354a2ef4ff79); the version 1.0 archive is included unchanged.
The study in brief
Learned routing controllers for wireless sensor networks are usually tied to the deployment they were trained on, because their value table or output layer is indexed by a fixed list of routes or node positions. The study defines the action-value function over per-candidate route features instead of action indices, so one network of 5121 parameters, trained with double deep Q-learning on a single topology distribution, scores route pools of 5 to 306 candidates without retraining. On 49 held-out topologies it improves delivery by 4.38 percentage points (distribution-free 95% interval [+1.72, +12.66]) and first node death by 272 packets [+79, +406] over a reactive reference. The advantage holds across nine structural regimes (+3.56 points, 319 topologies), and the learned policy maintains a measurable performance advantage across previously unseen topology distributions without retraining. A non-learned rule over the same features comes within 2.56 delivery points, while learning adds 148 packets of lifetime. Against an index-parameterized Q-network trained under the identical protocol, the controller leads by 4.26 delivery points [+1.38, +7.06] in distribution and by 3.48 points [+2.40, +4.66] under shift. All results are simulated.
Contents
· LPCAR_replication_package_v1.0.zip (unchanged): simulator/, data/, analysis/, tables/, protocol/, provenance/ and HASHES.md, described in its README.md.
· LPCAR_IDX_addendum.zip (new in v1.1): the simulator of the index-parameterized comparator (simulator/DRL7_IDX.m, derived from DRL7_RUN_ALL.m of version 1.0), its per-topology records on the transfer corpus and the four topology distributions, its protocol card and complete console log, the analysis scripts and the tables they regenerate, README_IDX.md and HASHES_IDX.md. Copy its folders into the root of the version 1.0 package; no existing file is replaced.
Terminology and numbering: the study calls the four node-placement processes (uniform, clustered, corridor, grid) topology distributions; file names and column labels keep the earlier label "family" so that every hash stays valid. Table numbers printed inside the version 1.0 files follow an earlier draft of the manuscript; the comparison with the index-parameterized network is Table 9 of the manuscript.
Reproducing
· The analysis, in seconds and without MATLAB: python3 analysis/run_all.py regenerates the results tables and the statistics quoted in the text from data/ and protocol/; analysis/idx_analysis.py and analysis/idx_table16.py regenerate the comparison with the index-parameterized network (Python 3 with NumPy, pandas, SciPy and Matplotlib).
· The simulation (MATLAB R2024a, several hours): DRL7_RUN_ALL('final') trains three networks and evaluates once on seeds 1056–1105; DRL7_RUN_ALL('baselines') loads the frozen networks and trains nothing; DRL7_IDX('idx') trains and evaluates the index-parameterized network under the same protocol. The final mode refuses to repeat the held-out evaluation once its stamp file exists. The trained networks are not in this archive, so baselines needs final (and, for the DQN rung, traindqn) to have run first, and DRL7_IDX needs the networks that final writes.
Provenance
The fingerprint 7cf77fea1129bedd is a deterministic hash over 50 configuration fields, including the network shape, learning rates, reward weights, state definition, the three seed lists and the weight-initialization seed. Recomputing it from the parameter values in provenance/DRL7_RUN_ALL_2026-08-15.m gives the fingerprint printed in the deposited protocol card. The comparator runs of version 1.1 reproduce the same fingerprint for the frozen configuration. This archive is a reproducibility and provenance record, not a pre-registration.
Disclosed deviations (detailed in the README files and in the manuscript)
· Five comparators (Min-hop, Linear, AOMDV, MMRE and DQN) were measured on the already-opened held-out corpus in evaluation-only passes that train nothing; they are exploratory.
· The QELAR reimplementation was corrected to its published coefficients after the held-out set was opened, as an evaluation-only re-run; the corrected version is the weaker of the two.
· AOMDV is implemented over node-disjoint route sets, whereas its authors specify link-disjoint routes.
· The evaluation log of version 1.0 prints a "REHEARSAL" banner from its evaluation-only branch; the banner does not describe the seeds, which are 1056–1105 as the protocol card records.
· The index-parameterized network inherits the hyperparameters chosen for the candidate-conditioned network and was not tuned separately. No policy was evaluated on the held-out seeds in version 1.1.
Scope
The energy model charges transmission and reception only (no idle listening, MAC or retransmission energy), the channel is lossless, and control-plane traffic is not charged. Absolute inference times are interpreted MATLAB on a laptop CPU; the study quotes their ratios and slopes, not the absolute microseconds.
License and citation
Code: MIT. Data and documentation: CC BY 4.0. Please cite the article (reference added on publication) together with this record. General geographic and genealogical information about the Teop language. Researchers from ILIMITED project partner Universidade de Aveiro (UAv) and colleagues from the Federal University of Rio Grande do Sul have co-authored the scientific article "Development of a COSMO-S… Researchers from ILIMITED project partner Universidade de Aveiro (UAv) and colleagues from the Federal University of Rio Grande do Sul have co-authored the scientific article "Development of a COSMO-SAC Parametrization with Advanced QM Method TZVPD-FINE" published in the journal Industrial & Engineering Chemistry Research.
The article highlights the development and validation of a new, high-accuracy parametrization for the COSMO-SAC model, a crucial tool in computational chemistry for estimating phase equilibria. Reproduction Package for the Paper “doi-checker: A Tool for Validating Bibliographies of Research Papers”
Abstract
This artifact is the reproduction package for the paper “doi-checke… Reproduction Package for the Paper “doi-checker: A Tool for Validating Bibliographies of Research Papers”
Abstract
This artifact is the reproduction package for the paper “doi-checker: A Tool for Validating Bibliographies of Research Papers”, accepted at ATIQSER 2026, co-located with ASE 2026.
The paper presents doi-checker, which extracts the references from a paper PDF and verifies the DOIs they print against Crossref and DataCite, and compares it with four other open-source reference checkers over 3888 papers published at CAV, FSE, ISSTA, PLDI, POPL and TACAS between 2020 and 2026.
The results here are those of doi-checker 1.0, which reads the numbered [N] reference list a PDF prints itself — via pdftohtml, and taking DOIs from the PDF’s own hyperlinks — and falls back to GROBID for bibliographies of any other shape. That covers 2275 of the 3888 papers; the remaining 1613 go through GROBID as before.
The artifact contains the five tools at the commits that produced the results, the caches and database mirrors that make a corpus-scale run feasible, the analysis that turns raw results into every number in the paper, and the original results of the paper — so the analysis can be checked without re-running anything.
It also contains 2500 of the 3888 papers: every one whose licence permits redistribution. The rest cannot legally be bundled, and make corpus fetches them from the publishers. See The corpus and its licences — this affects what a re-run can cover, and nothing else.
Everything installs and runs with no network at all. There are three ways to run the experiments, cheapest first: a smoke test over 3 papers, a subset of 949 papers (~1 h at JOBS=16), and the 2500 bundled papers (~6 h).
Tested in the VM
Run end to end from this archive unzipped, on a freshly imported, unmodified CAV 26 VM whose network adapter had been removed entirely (--nic1 none), so nothing could reach the internet even by accident. Every entry in the archive except this file is byte-identical to the one those runs were made from:
Step
Result
Time
./setup_vm.sh
podman installed offline, GROBID image loaded, all checks pass
63 s
./install.sh
5 virtualenvs built from bundled wheels, all four caches re-keyed to 100% coverage of the bundled corpus
58 s
make smoketest OFFLINE=1
overall: PASS; every reference count matches the paper, and doi-checker matches it in full — all 15 references of the paper it reads with pdftohtml, identical in text, DOI and verdict
2 min 44 s, or 32 s with NO_GROBID=1
make compare-paper RESULTS_DIR=experiments/results-paper
identical to the paper — all 1280 macros the paper typesets regenerate byte-for-byte
17 s
make check-corpus
2500 present, 1388 not bundled, 0 missing
1 s
Those times are for a host whose page cache is still warm from unpacking the archive. Cold, the same steps took 4 min 47 s and 3 min 23 s — it is the shared folder being read a page at a time, not the work itself, and it happens once.
The one check that cannot be run from the archive is re-cutting the corpus out of the publishers’ volumes, because those volumes are not bundled — 5 GB, and needed only to redo the split. It was verified in the development repository instead: re-cutting the 65 CAV 2020 papers reproduced all 65 byte-identically in 46 s. From the package, make proceedings fetches the volumes first; see Fetching the rest.
Where to find things
Requirements
disk, memory, and why the artifact does not fit inside the VM
TL;DR
the whole reproduction path, copy-pasteable
What to Check, and How
the paper’s five claims, and the command for each
Contents
what is in the package, and how big
Experiments
the three benchmark sets
Smoke Test
exactly what it should print
Results
comparing your run with the paper’s
Running Offline
what makes this work without a network, and the one thing it costs
The corpus and its licences
why 2500 of 3888 papers ship, and how to get the rest
Information for Reuse
other papers, and adding a sixth tool
Provenance and Licences
commits and terms
Rebuilding this Package
the reproducible build
Known Issues
what can go wrong, and what it means
Requirements
The tools are compiled for x86_64 and will not run on ARM. Any x86_64 Linux with Python 3.11+ works; the CAV 26 VM (Ubuntu 24.04.4, user and password cav, Python 3.12.3) is what it was tested on.
Disk
9.5 GB archive, 21.1 GB unpacked, ~0.6 GB for a full run’s results
Memory
4 GB; 8 GB is comfortable if you also run GROBID
Cores
any — this work is network- and I/O-bound, not CPU-bound
Network
none, for everything in this README
The CAV 26 VM’s own disk is 25 GB with about 12 GB free, so the artifact does not fit inside it. Unpack it on the host and mount it into the VM as a shared folder, as in the TL;DR below. That is the route that was tested.
doi-checker 1.0 shells out to pdftohtml (from poppler-utils) for its default extractor. It is bundled in vendor/deb/ and installed by setup_vm.sh along with everything else, so there is nothing to fetch. Without it the tool still runs — every paper then falls back to GROBID, as in the pre-1.0 releases — but the DOI coverage would drop back to the pre-1.0 figures.
The bundled pdftohtml is 24.02 (what Ubuntu 24.04 ships, so that setup_vm.sh installs cleanly on the VM), while the shipped results were produced with 26.01 on the build machine. The two do not emit the same XML: on a 60-paper sample, drawn from the papers 1.0 reads with this extractor, not one of the 60 XML documents matched byte for byte. The reference lists parsed out of them did — all 60, same reference text and same DOIs. So the version difference does not move a single number in the paper; it only means that a cache regenerated on one version is not byte-identical to one regenerated on the other.
That holds only while the shipped XML is the XML being read, which is what install.sh arranges by touching the cache (see Known Issues — a reproducible archive unpacks the cache and the PDF with the same timestamp, and doi-checker requires the cache to be strictly newer). Skip that step and the tool re-runs pdftohtml for every paper, and the 60-paper agreement above is the only evidence that the local poppler would have agreed. So: run install.sh, and check rekey_caches.py --check reads 100% for the pdftohtml XML line before concluding anything from a run that disagrees with the paper.
The VM ships build-essential, git, curl, make, python3-full, Java and Rust, but no container engine, which GROBID needs. setup_vm.sh installs podman and the shared libraries the prebuilt hallucinator binary needs from the bundled vendor/ directory, then loads the GROBID image — all without a network. It uses sudo for apt and nothing else. Nothing is installed system-wide beyond those packages: each Python tool gets a private virtualenv (the Rust one ships as a binary), and every cache, database and result stays inside the artifact directory.
If podman cannot be installed, add NO_GROBID=1 to any run target: the bundled extraction caches cover every paper in the corpus, so nothing needs re-extracting. Only checking a new PDF needs GROBID — and since 1.0, only one whose bibliography is not a numbered [N] list, because pdftohtml reads those locally with no service at all.
TL;DR
Step 0, on the host. Unpack the archive and give the VM the directory as a shared folder, with the VM shut down:
unzip doi-checker-artifact.zip -d ~/cav-artifact # 21.1 GB, on the host
VBoxManage import cav2026.ova # if you have not already
VBoxManage sharedfolder add cav2026 --name artifact \
--hostpath ~/cav-artifact --automount
# Not optional: every Python virtualenv contains a symlink, and vboxsf refuses
# to create one unless the host says it may. There is no equivalent setting in
# the VirtualBox GUI. Without this, ./install.sh stops with
# Error: [Errno 1] Operation not permitted: 'lib' -> '.../.venv/lib64'
VBoxManage setextradata cav2026 \
VBoxInternal2/SharedFoldersEnableSymlinksCreate/artifact 1
# The 1 vCPU / 4 GB default works but is slow, and GROBID wants headroom.
VBoxManage modifyvm cav2026 --memory 8192 --cpus 4
VBoxManage startvm cav2026
Steps 1–4, inside the VM, in a terminal at /media/sf_artifact/doi-checker-artifact. Every command in this README is run from there.
# 1. Set up. setup_vm.sh asks for the sudo password once; on this VM it is `cav`.
./setup_vm.sh
# -> VM is ready. Next: ./install.sh
./install.sh
# -> ... refchecker bibliographies 2499 / 2500 (100.0%)
# 2. Check the analysis reproduces the paper, without running any experiment.
# 25 seconds, and the most direct check in the artifact.
make compare-paper RESULTS_DIR=experiments/results-paper
# -> identical to the paper
# 3. Check the pipeline itself works, on three papers.
# ~40 min the first time (nothing has been read off the 19 GB yet), ~5 min after.
make smoketest OFFLINE=1
# -> overall: PASS
# 4. Optionally re-run more of the experiment, then compare again.
make subset-experiments OFFLINE=1 # 949 papers, ~1 h at JOBS=16
make data-analysis
make compare-paper
make help lists every target.
If you cannot use a shared folder
Give the VM a bigger disk. VirtualBox cannot resize a VMDK, so clone it to VDI first (VBoxManage clonemedium), resize that, attach it, and grow the partition inside the VM with growpart and resize2fs.
Make the artifact smaller. Deleting the two database mirrors brings it to 7.4 GB, which fits — but it changes the results, so read Running on a small disk first.
What to Check, and How
The paper asks five research questions. All five are answered by macros in data-analysis/output/macros.tex, which is the single generated file the paper \inputs — it hard-codes no figure of its own, so every number and every table cell in the evaluation comes from that file. make compare-paper therefore checks all five claims at once.
They do not all reproduce equally well from an offline run, and it is worth knowing which is which before starting:
Paper
Question
Comes from
Offline?
RQ1 Coverage
what fraction of references print a DOI, by venue and year
the doi column of references.csv
yes — extraction only
RQ2 Validity
how many printed DOIs are invalid or non-resolvable
doi-checker’s invalid_doi flags
no — see the one difference
RQ3 Integrity
how many DOIs resolve but disagree with the cited entry
doi-checker’s author_mismatch / title_mismatch flags
yes — from cached registry records
RQ4 Accuracy
how often doi-checker’s own verdict is right
the hand-labelled sample, not a run
yes — an input, not an output, but see the note below
RQ5 Comparison
how doi-checker compares with the other tools
all three tools’ results, joined on printed DOIs
yes, to within the differences below
Three checks, in increasing cost:
Check
Command
Time
Proves
The analysis reproduces
make compare-paper RESULTS_DIR=experiments/results-paper
25 s
every number in the paper follows from the shipped results
The pipeline works
make smoketest OFFLINE=1
~3.5 min
the tools run and agree with the paper on 3 papers
The runs reproduce
make full-experiments OFFLINE=1 then make compare-paper
~6 h
the results themselves, over the 2500 bundled papers
Those are different claims that fail for different reasons, which is why they are separate.
RQ4’s sample was drawn against an earlier version, and has been re-keyed onto 1.0. The 200 references were drawn once and 53 labelled by hand, and a row identifies a reference by its position in the paper’s reference list — which 1.0 renumbers for the 2275 papers it reads from their own numbered list. The rows are therefore matched onto this run by reference text: all 200 were found again, and the hand labels carried over, because a label judges a reference and it is still the same reference. Where 1.0’s verdict on a labelled reference changed, the label was remapped by what it already asserts — an extraction label whose DOI is now recovered becomes correct, an incorrect label whose false report is gone becomes correct, and a confirmed invalid report that 1.0 no longer makes becomes declined, a label added for it. Three references that none of those rules covered were judged by hand against the registry record.
Two consequences are worth knowing before reading the RQ4 figures. The strata were equal-sized when drawn and are not any more — 31 of the 53 labelled references now sit in one category — so RqAccuracyWeighted, RqPrecision and the per-category rates rest on very few references each. And declined is now the largest error class at 11 of 53: those references print an unregistered 10.5555 identifier as a DOI, and 1.0 passes them over as needing no DOI at all. make rq-sample draws a fresh, unlabelled sample if you would rather re-label from scratch; data-analysis/output-paper/rq4_sample.csv is the re-keyed one these figures come from.
Contents
Path
What it is
setup_vm.sh
Installs podman, the GROBID image and the missing libraries, all from vendor/. --verify only checks.
install.sh
Builds a virtualenv per Python tool from the bundled wheels, then re-keys the caches to wherever this was unpacked.
Makefile
Every target used in this README; make help lists them all.
data/<venue>/<year>/
The corpus: the 2500 published papers this artifact may redistribute, one PDF each, with a papers.json manifest per venue-year covering all 3888 and giving each paper’s title, DOI, page range and file name.
data/proceedings/
How the corpus was made and what may be shared: fetch_proceedings.py, split_proceedings.py, check_corpus.py, volumes.yml, and licences.csv — the per-paper licence audit.
tools/tools.yml
Each tool under comparison, pinned to a commit (doi-checker to its 1.0 tag), and how it is built and invoked.
tools/<name>/
src (the clone at that commit), COMMIT, venv or bin, and cache.
tools/data/
The offline mirrors: DBLP, ACL, arXiv and IACR as SQLite for hallucinator; DBLP, ACL and a seeded Crossref in refchecker’s schema.
experiments/
The runner (run_experiments.py), the per-tool adapters (adapters.py), and the paper lists.
experiments/results-paper/
The raw results the paper reports — 11 664 per-paper JSON records, the two CSV tables, and tools.json pinning each tool’s commit.
data-analysis/
Results → the LaTeX macros the paper typesets, plus rq4_analysis.csv, the hand-labelled sample behind RQ4.
data-analysis/output-paper/
What that analysis produced for the paper: macros.tex, tool_summary.csv, formats.json, environment.json, RQ4 material.
vendor/
What the install needs and the VM lacks: Ubuntu .debs as an apt repository, a CPython 3.12 wheelhouse per virtualenv, the GROBID image. MANIFEST.sha256 checksums all of it; LICENSES.md and licences/ carry every component’s licence and its full text.
packaging/
Pinned dependency versions, the offline installer, and the cache re-keying.
LICENSE
Terms for the artifact and for everything bundled with it.
experiments/results/ and data-analysis/output/ do not exist until you run something — deliberately, so a fresh run cannot be confused with the shipped one.
Where the 21.1 GB goes:
Size
Files
tools/data/refchecker/ — refchecker’s mirrors
8.54 GB
3
tools/data/ — hallucinator’s mirrors
5.25 GB
5
data/ — the bundled corpus
2.71 GB
2541
tools/*/cache/ — query and extraction caches
3.00 GB
216 035
vendor/ — packages, wheels, GROBID image
0.81 GB
267
experiments/results-paper/
0.58 GB
11 667
tools/*/src, tools/hallucinator/bin
0.09 GB
4046
code, documentation, analysis outputs
2 MB
45
The experiments cover 3888 papers — CAV 498, FSE 1199, ISSTA 600, PLDI 600, POPL 531, TACAS 460, spanning 2020–2026 except ISSTA which ran to 2025 — of which the 2500 that may be redistributed are bundled; see The corpus and its licences.
Experiments
The three benchmark sets
All are resumable at (tool, paper) granularity — a pair whose result file exists is skipped — so Ctrl-C and re-run costs only what was in flight, and a subset run is a head start on the full one. That is measured, not asserted: interrupting a subset run after four minutes left
WARNING: interrupted: killed 16 running tool(s); 83 of 2847 pair(s) finished
and kept -- re-run to continue from here
with papers.csv and references.csv written for those 83, and the next run opened with 83 (tool, paper) pair(s) already done; skipping and carried on from there. Interrupting while GROBID is still starting is also safe — it says so and exits, having run nothing.
Set
Papers
Command
Cost
Smoke test
3
make smoketest OFFLINE=1
~3.5 min (~30 s with NO_GROBID=1)
Subset
949
make subset-experiments OFFLINE=1
~1 h at JOBS=16
Bundled corpus
2500
make full-experiments OFFLINE=1
~6 h at JOBS=16
The subset spans all 41 venue-years — 4 to 26 papers from each, median 25 — drawn round-robin so no venue drops out and cheapest-first within each so the budget buys as many papers as possible. data-analysis/create_subset.py sized it from the measured per-paper cost of the paper’s own run — 16.0 hours of work, one hour of wall clock at JOBS=16 — and it ships as experiments/subset.txt, so it is the same 949 papers for everyone. make subset HOURS=4 re-sizes it. It is drawn only from the papers that ship, so it runs out of the box.
It is enough to see every venue, every bibliography style and every tool behave as the paper describes. It is not enough to reproduce the paper’s numbers: cheapest-first selection biases it towards short papers with few references, so its rates are not the corpus’s rates.
The bundled corpus is the 2500 papers this artifact may redistribute × 3 tools. The paper’s own run covered all 3888 and cost 147 hours of work, about 9 hours of wall clock at JOBS=16; two thirds of that is about 6 hours. Offline it is faster. To run over all 3888, fetch the rest first with make corpus.
Running them
JOBS is the number of papers in flight and is what sets the wall clock. It is not the core count: 16 concurrent papers kept 0.65 of 16 cores busy — 96% idle — so JOBS=16 is reasonable even on a 4-core VM. Every knob is documented at the top of the Makefile:
make full-experiments OFFLINE=1 JOBS=32 WORKERS=16 GROBID_REPLICAS=3 # bigger machine
make full-experiments OFFLINE=1 TOOLS="doi-checker" # one tool
make full-experiments OFFLINE=1 NO_GROBID=1 # no container engine
A run writes one JSON record per tool and paper under experiments/results/<tool>/<venue>/<year>/<paper>.json, and at the end results/references.csv (a row per reference) and results/papers.csv (a row per paper).
Two of the five tools, halref and verifyref, are excluded by default. Both are bundled and both run, but each needs 340–900 s per paper against roughly 9 hours for the other three between them, so on the full corpus they are ~150 h apiece — a cost the paper reports rather than a run it makes:
make full-experiments TOOLS="halref verifyref" # note: no OFFLINE=1
Do not pass OFFLINE=1 to those two. Neither has a cache or an offline mode of any kind, so offline they would not run slowly — they would fail on every reference and produce a result set that says more about the proxy than about the tools. They are the one part of this artifact that genuinely needs a network.
Smoke Test
make smoketest OFFLINE=1
Three deliberately cheap papers, one per bibliography style in the corpus, 59 references between them, three tools over each. All three are CC-BY, so all three ship with the artifact. Add NO_GROBID=1 if setup_vm.sh could not give you a working container engine — the results are the same and it is much faster.
Everything it prints is captured, so “did it work” is answered by reading one file rather than by watching a terminal:
experiments/results/smoketest.log everything it printed
experiments/results/smoketest_summary.txt a short, diffable summary
Measured in the offline VM:
make smoketest OFFLINE=1
3 min 20 s
make smoketest OFFLINE=1 NO_GROBID=1
31 s
Almost all of that is GROBID starting up — 2 min 21 s of it, measured, before the first paper is touched. It then never has to answer a question, because the extraction caches are complete, which is why NO_GROBID=1 gives the same results in a fraction of the time. If the run seems to hang after starting GROBID, it is loading its models; give it three minutes.
It can be much slower the first time, if the caches and mirrors have to come off the host’s physical disk rather than its page cache: the tools read them through the shared folder a page at a time. Unpacking the archive immediately before leaves the host’s cache warm and the first run then costs about the same as any other — 3 min 20 s, measured. Come back to a cold machine and the same run has taken 37 minutes. Either way it is a one-off, and it is I/O, not the tools.
Expected result
It ends with overall: PASS and a line per tool:
=== smoke-test summary ===
PASS doi-checker @5b40c64048b8 papers=3/3 refs=59 checked=38 flagged=0
PASS hallucinator @4c3821695a9c papers=3/3 refs=35 checked=29 flagged=7
PASS refchecker @7673d0189f26 papers=3/3 refs=58 checked=51-54 flagged=24-28
overall: PASS
The commit after each @ is the exact build of that tool, and the same one recorded in experiments/results-paper/tools.json.
Per paper, beside what the paper’s own online run recorded for the same three:
Tool
Paper
Refs
Checked
Flagged
The paper
doi-checker
fse/2023/0191_p4b-…
15
10
0
same
doi-checker
tacas/2024/0004_speculative-…
13
0
0
same
doi-checker
popl/2024/0002_deciding-…
31
28
0
same
hallucinator
fse/2023/0191_p4b-…
15
12
2
flagged 1
hallucinator
tacas/2024/0004_speculative-…
13
12
2
flagged 1
hallucinator
popl/2024/0002_deciding-…
7
5
3
same
refchecker
fse/2023/0191_p4b-…
14
11–12
7
same
refchecker
tacas/2024/0004_speculative-…
13
10–12
4–5
flagged 5
refchecker
popl/2024/0002_deciding-…
31
30
14
same
Every reference count reproduces the paper exactly, and doi-checker reproduces in full — same references extracted, checked and flagged, on all three papers. hallucinator flags two references offline that it resolved online, which is what the offline caveat predicts.
The three papers also exercise both extractors, which is why they were chosen: the FSE paper’s [N] list is read by pdftohtml, while the TACAS (LNCS) and POPL (author–year) bibliographies are not numbered [N] lists and fall back to GROBID. The log names the extractor used for each paper.
refchecker is the one tool whose numbers move between runs even on the same machine: it verifies a paper’s references through a thread pool and does not always finish the same ones before deciding it is done. Six runs in the VM gave checked between 51 and 54 against the paper’s 54, and flagged between 24 and 28. Treat its checked and flagged as a range, not a fingerprint; the other two tools are stable.
Three entries in that table are results the paper reports, not failures:
hallucinator finds 7 references in the POPL paper where the others find 31. It segments bibliographies on [1] and 1. markers, and the unnumbered ACM author-year style used by POPL and PLDI has neither.
doi-checker checks 0 of the TACAS paper’s 13 references. It only judges references carrying a DOI, and that LNCS bibliography prints none. Out of scope is not the same as clean, so those are recorded as unverified.
refchecker flags far more than the others. It counts citation-quality findings — a published paper cited as an arXiv preprint, a missing URL — that no other tool looks for. The analysis reports a strict and a broad rate separately for exactly this reason.
The smoke test re-runs those three papers even if results exist (--force), and caps each tool at 300 s rather than the 900 the real runs use: it asks whether the pipeline works, not whether a slow tool can finish a large bibliography.
Results
make data-analysis # results/ -> data-analysis/output/macros.tex
make compare-paper # and diff those macros against the paper's
The paper’s own results ship beside the folders a run creates:
Created by a run
Shipped with the paper
experiments/results/
experiments/results-paper/
data-analysis/output/
data-analysis/output-paper/
So the analysis can be checked on its own, without spending hours on the runs:
make compare-paper RESULTS_DIR=experiments/results-paper
# -> identical to the paper
Over a smoke test or a subset that same diff is large and uninformative — those are not the corpus, so their rates are not the corpus’s rates. It is meaningful after a full run, and exact against experiments/results-paper. data-analysis/output/tool_summary.csv is the readable one-line-per-tool summary.
Three inputs the analysis needs are not produced by any run: the hand-labelled RQ4 sample, the machine description, and the per-venue page formats. They ship under output-paper/ and the Makefile falls back to them, which is why the comparison comes out identical without your regenerating anything. If you do regenerate them (make environment, make formats), yours are used instead and the environment macros will legitimately differ — they describe your machine, not the one in the paper.
RQ4 measures how often the tool is right, which no run can produce: 200 references were drawn once and labelled by hand, and the paper reports the 53 that carry a label. The sample ships twice, byte-identically — as data-analysis/rq4_analysis.csv (an input, beside the code) and as data-analysis/output-paper/rq4_sample.csv (where the analysis looks). make rq-sample draws a fresh sample and refuses to overwrite an existing one; a fresh sample carries no labels, so its RQ4 macros are zero until somebody labels it.
Running Offline
OFFLINE=1 does not merely assume there is no network — it enforces one, by pointing the tools’ proxy variables at a closed port (localhost is exempt, so GROBID still works). An uncached lookup therefore fails visibly instead of quietly topping the cache up from the live APIs, which is the only way “the bundled caches are sufficient” can be tested rather than assumed.
Cache
Size
Covers
tools/doi-checker/cache/
1.90 GB
pdftohtml XML and GROBID TEI for 100% of the corpus, plus every Crossref/DataCite record that resolved
tools/refchecker/cache/
1.08 GB
parsed bibliography for 100% of the corpus, plus its API responses
tools/hallucinator/cache/queries.sqlite
25 MB
its query cache across the whole corpus
tools/data/*.sqlite
5.25 GB
hallucinator’s DBLP, ACL, arXiv and IACR mirrors
tools/data/refchecker/*.db
8.54 GB
refchecker’s own DBLP, ACL and Crossref mirrors
The first three make a re-run of this corpus offline. The mirrors are what let the tools answer questions about papers that are not in it, so they matter for reuse rather than for reproduction.
1.7 GB of doi-checker’s cache is the pdftohtml XML of the 3879 papers it could read, and that part is the one cache in the artifact that needs no network and no service to rebuild — pdftohtml is a local binary. Deleting tools/doi-checker/cache/*-pdftohtml.xml reclaims it, and the next run regenerates only what it touches, at the same verdicts (see the version note under Requirements: the regenerated XML differs from the shipped XML, the references parsed out of it do not). The 218 MB of GROBID TEI beside it is not replaceable that way: without a container engine it can only be read, not remade.
Leaving OFFLINE=1 off
Everything also runs online, and nothing needs an API key. Measured in the same VM with networking enabled and no .env:
Offline
Online
make smoketest
3 min 20 s, PASS
6 min 19 s, PASS
doi-checker
59 refs, 38 checked, 0 flagged
identical
hallucinator
35 refs, 29 checked, 7 flagged
identical
refchecker
58 refs, 53 checked, 26 flagged
55 checked, 26 flagged
So online is slower and, on these three papers, reaches the same verdicts. The log carries two warnings that are expected rather than faults:
No Crossref --mailto given — the anonymous pool has lower rate limits. Copy .env.example to .env and set CROSSREF_MAILTO to move into the polite one.
rate limit notices from the tools’ own retry logic, for the same reason.
Online is what to use for the one thing the caches cannot answer, below. For everything else OFFLINE=1 is faster, and it is the only way to be sure the result came from the bundled data rather than from today’s APIs.
The one difference, offline
doi-checker reports no invalid DOIs when run offline. It caches a lookup only when the DOI resolved; an unregistered DOI is re-checked every run, by design, because “unknown to Crossref and DataCite” is a claim it will not make from a stale cache. Offline it cannot tell an unregistered DOI from an unreachable registry, so it draws no conclusion. Over the full corpus:
doi-checker
Online
Offline
invalid_doi
452
0
author_mismatch
2748
2748
title_mismatch
3951
3951
references flagged
5924
5472
5472 exactly, not approximately: every one of the 452 carries invalid_doi and nothing else, so losing that flag unflags the reference entirely. The other two flag types come from cached registry records and are unaffected, as are hallucinator’s and refchecker’s results. This is RQ2, and only RQ2.
To reproduce those numbers exactly, run that one tool with a network — it is the cheapest of the three at a few seconds per paper, and the corpus prints 43 741 distinct DOIs, so that is how many registry lookups it makes — a couple of hours rather than a day. Note that Crossref rate-limits an anonymous IP to 10 requests a second, so JOBS well above 8 buys 429s rather than throughput on this one target:
make full-experiments TOOLS="doi-checker" # note: no OFFLINE=1
Running on a small disk
The two mirror sets are 13.7 of the artifact’s 21.1 GB, and the tools detect their absence and fall back:
rm -rf tools/data/refchecker tools/data/arxiv.sqlite tools/data/dblp.sqlite
That leaves 7.4 GB, which fits inside the VM’s own disk, and everything still installs and runs offline — but it changes the answers, and not subtly. Measured: the same smoke test, in the same VM, with and without them.
With the mirrors
Without
wall clock (NO_GROBID=1)
31–54 s
4 min 55 s
doi-checker
59 refs, 38 checked, 0 flagged
unchanged — it does not use them
hallucinator
3/3 papers, 35 refs, 7 flagged
2/3 papers, 20 refs, 17 flagged
refchecker
58 refs, 24–27 flagged
58 refs, 41 flagged
Every lookup a mirror would have answered now cannot be answered at all, so the reference is reported as not_found and counted as a finding. hallucinator does not even finish one of the three papers, and the summary still says PASS — it checks that the tools ran, not that they were well supplied.
So this is a way to make the artifact fit, not a way to reproduce the paper on a smaller disk. If disk is the constraint, prefer the shared folder.
The corpus and its licences
The experiments were run over 3888 papers. 2500 of them ship with this artifact — every paper whose licence permits a third party to redistribute it. The other 1388 cannot legally be bundled, and are fetched from the publishers by make corpus.
Papers
Ships?
CC-BY 4.0
2430
yes
CC-BY-SA 4.0
45
yes
CC-BY-ND 4.0
25
yes — redistribution of the unmodified work is permitted, and these are bundled unmodified
CC-BY-NC, NC-SA, NC-ND
193
no — would impose a NonCommercial condition on everyone reusing this artifact
ACM default terms
1195
no — all rights reserved; the notice permits personal and classroom copying, not redistribution
Which paper falls where, and on what evidence, is recorded per paper in data/proceedings/licences.csv. Each paper was checked against two independent sources — the per-article licence Crossref holds, and the notice printed in the PDF itself — which agree on 2314 papers, with one source silent on the rest and no outright conflicts. Licensing is per paper, not per proceedings volume: an open-access article sits in the same volume as an all-rights-reserved one, which is why this is a per-paper filter and not a per-venue one.
Bundled papers are byte-identical to the publisher’s PDF, so the attribution CC-BY requires travels with each file as its own printed notice.
What ships, by venue
Venue
Ships
Of
Note
CAV
498
498
CC-BY throughout
TACAS
400
460
TACAS 2026 moved to CC-BY-NC-ND — 60 of its 72 chapters
POPL
497
531
FSE
515
1199
PLDI
416
600
ISSTA
174
600
ACM’s move to open access is the whole story on their side: 24% of their papers were redistributable in 2020, 49% in 2023, 90% in 2026.
What this changes, and what it does not
It does not affect the paper’s results. experiments/results-paper/ covers all 3888 papers, so make compare-paper — the headline check — is untouched, and so are the smoke test and the subset, both of which name only bundled papers.
What it limits is a fresh full run: out of the box that covers the 2500 bundled papers rather than all 3888. make check-corpus tells you exactly where you stand, and distinguishes the two reasons a paper can be absent:
make check-corpus # present / not bundled / missing, per venue-year
make check-corpus --full # treat "not bundled" as missing too
“Not bundled” is expected and needs no action. “Missing” means something is wrong — a broken unpack, or a rebuild that came out wrong.
Fetching the rest
make corpus # download the publishers' volumes, split, check
This is the one part of the artifact that needs a network, and it needs rather more besides:
Network
yes — it downloads from ACM and Springer
Browser
a Chromium-family browser on a real X display. Both publishers serve these PDFs only to a real browser and refuse scripted clients, so the fetcher drives one over the DevTools Protocol rather than pretending to be one. Set $CHROME if yours is under an unusual name. Run it from the VM’s desktop session, not over ssh.
Tools
qpdf (required, to merge ACM’s per-article PDFs), pdftk (optional, for bookmarks)
Disk
~40 GB for the publishers’ volumes, plus 4.8 GB for the split papers
Time
hours
The 40 GB does not fit on the CAV VM’s own disk, so point SRC_DIR at the shared folder or run the rebuild on the host:
sudo apt-get install -y chromium-browser qpdf pdftk
make corpus SRC_DIR=/media/sf_artifact/proceedings
make proceedings-list shows every volume and whether it is present, and needs nothing installed. Both stages are resumable and skip what is already there.
The download half needs a desktop session. Run inside the VM over ssh or VBoxManage guestcontrol it stops with
no DISPLAY: the publishers' bot checks fail in headless Chrome, so this needs
a real X display.
which is accurate rather than a limitation to work around: xvfb-run does not help, because the checks detect it. Open a terminal in the VM’s desktop, or do the download on the host.
The split half needs neither a network nor a browser, so once the volumes are there, make split re-cuts every one of them, and make split-one re-cuts a single volume:
make split # all of them
make split-one PDFS="data/proceedings/cav20-part*.pdf" OUT=data/cav/2020 # just one
Both need data/proceedings/*.pdf, which the archive does not carry, so from an unpacked package they stop with no proceedings PDFs found in data/proceedings until make proceedings has fetched them.
That the split is deterministic was verified in the development repository, where the volumes live: re-cutting the two CAV 2020 volumes produced all 65 papers byte-identically — same file names, same SHA256 — in 46 seconds. That is the property the caches depend on, since they are keyed on each paper’s file name.
The manifests this is all checked against — data/<venue>/<year>/papers.json — always ship, whatever the licence, and record each paper’s title, DOI, page range and the file name it was split to. That last part is what makes a rebuild verifiable rather than merely plausible: the tool caches are keyed on each paper’s file name, so a corpus that splits under different names is a silent cache miss, not a visible error.
Information for Reuse
Nothing here is wired to the corpus it ships with.
Check a PDF that is not in the corpus. Any of the five tools runs directly; the ones that extract through GROBID need it running:
podman run -d --rm --name grobid -p 8070:8070 docker.io/lfoppiano/grobid:0.8.1
tools/doi-checker/venv/bin/doi-checker --json \
--grobid-url http://localhost:8070 \
--cache-dir tools/doi-checker/cache \
--title-similarity 80 /path/to/your.pdf
podman rm -f grobid
Since 1.0 the container is only the fallback. If the PDF prints a numbered [N] bibliography, --extractor xml reads it with pdftohtml alone and needs no GROBID and no network beyond the registry lookups:
tools/doi-checker/venv/bin/doi-checker --extractor xml \
--cache-dir tools/doi-checker/cache \
--title-similarity 80 /path/to/your.pdf
--extractor grobid forces the old behaviour, which is what to use to compare against the pre-1.0 numbers.
Run the whole comparison over your own papers. Drop them under data/<venue>/<year>/ and hand the runner a list; venue and year are read off the path and nothing else has to be registered:
printf 'data/mine/2026/paper1.pdf\ndata/mine/2026/paper2.pdf\n' > mine.txt
.venv/bin/python experiments/run_experiments.py --paper-list mine.txt \
--tools doi-checker hallucinator refchecker --jobs 8
Leave OFFLINE=1 off for papers outside the corpus: the bundled query caches know nothing about their references. The mirrors in tools/data/ still help, since they cover DBLP, ACL, arXiv, IACR and a seeded Crossref regardless of which papers cite them.
Add a sixth tool. One entry in tools/tools.yml saying how to fetch and build it, and one adapter in experiments/adapters.py saying how to invoke it and how to map its output onto the shared ok / flagged / unverified vocabulary. The adapter always stores the tool’s own output beside the normalised record, so a mapping that turns out to be wrong can be corrected without re-running anything. tools/README.md documents both, and the five existing adapters are worked examples of five quite different output formats.
Take just the mirrors. tools/data/ is a general-purpose offline bibliographic dataset — DBLP, the ACL Anthology, arXiv and IACR ePrint as SQLite, plus refchecker’s own schema — useful well outside this paper. The scripts that build them are in tools/setup_tools.py, driven by the databases: section of tools.yml.
Provenance and Licences
Each tool is pinned to an exact commit, recorded both in tools/<name>/COMMIT and in experiments/results-paper/tools.json. Those agree, so the bundled sources are the ones that produced the shipped results.
Tool
Commit
Licence
Upstream
doi-checker
5b40c64048b8 (tag 1.0)
Apache-2.0
https://gitlab.com/sosy-lab/software/doi-checker
hallucinator
4c3821695a9c
GPL-3.0
https://github.com/gianlucasb/hallucinator
halref
64903c127891
MIT
https://github.com/davidjurgens/hallucinated-reference-finder
refchecker
7673d0189f26
MIT
https://github.com/markrussinovich/refchecker
verifyref
1a4d3e4ebe89
GPL-3.0
https://github.com/hadipourh/verifyref
GROBID (image)
0.8.1
Apache-2.0
docker.io/lfoppiano/grobid:0.8.1
The code written for this artifact — the runner, the adapters, the analysis and the packaging — is Apache-2.0; see LICENSE, which also states what the bundled third-party material is and under what terms it travels.
The vendored binaries — 93 Ubuntu packages, 108 Python wheels and the GROBID image — are catalogued in vendor/LICENSES.md, which gives each component’s version, the licence it declares, and a link to its full text under vendor/licences/. Every Ubuntu package’s Debian copyright file is reproduced there in full, including the ones Debian shares between packages built from a single source; a handful of wheels ship no licence file at all and are marked as such. Several of those packages are GPL or LGPL, so that file also records where their corresponding source is: each .deb is an unmodified, version-pinned binary from the Ubuntu 24.04 archive, and nothing in vendor/ was rebuilt or patched here.
The paper PDFs in data/ are the published versions, split out of the publishers’ proceedings volumes. Copyright in them remains with the respective authors, and nothing here relicenses them: each is redistributed under the Creative Commons licence its publisher applied to it — CC-BY, CC-BY-SA or CC-BY-ND — recorded per paper in data/proceedings/licences.csv. They are bundled unmodified, so the attribution those licences require travels with each file as its own printed notice. Papers under any other terms are not included; see The corpus and its licences.
Rebuilding this Package
For anyone extending the work rather than evaluating it, from the development repository:
make vendor # the only step that needs a network
make package # dist/doi-checker-artifact.zip + .sha256
make package-verify # build it twice and check the two are byte-identical
The build takes about four minutes and is reproducible: entries are sorted by archive path, timestamps fixed to SOURCE_DATE_EPOCH (defaulting to the last commit) and permissions normalised, so the same repository and the same vendor/ give the same archive byte for byte — and therefore the same SHA256 as the copy distributed with this package. vendor/MANIFEST.sha256 pins the downloaded inputs, the only part a network can change. Already-compressed members (PDFs, wheels, .debs, the image) are stored rather than deflated, which is why a 9.5 GB archive builds in minutes rather than hours.
make package CORPUS=0 leaves the paper PDFs out (keeping the manifests, so make corpus can rebuild them); MIRRORS=0 leaves out the database snapshots.
Known Issues
install.sh stops with Operation not permitted: 'lib' -> '.../lib64'. The artifact is on a VirtualBox shared folder, which refuses to create the symlink every Python virtualenv contains. Shut the VM down and run the VBoxManage setextradata … SharedFoldersEnableSymlinksCreate command from the TL;DR on the host, then start it again. install.sh checks for this before doing anything else and prints the command with your share’s name filled in.
sudo asks for a password. On the CAV VM the cav user’s password is also cav. Only setup_vm.sh needs it, and only for apt. To run it with no terminal at all, set SUDO="sudo -A" and SUDO_ASKPASS to a script that echoes the password.
podman info fails, or podman will not install. Run everything with NO_GROBID=1. The bundled extraction caches cover the corpus, so nothing in this README is lost; only checking a new PDF needs GROBID.
doi-checker reports ACM’s 10.5555 prefix inconsistently between its two extractors. Of the corpus references that print doi.org/10.5555/… — an identifier the ACM Digital Library assigns but does not register — 1.0 flags 94 as an invalid DOI and passes over 223 as needing no DOI, and the split is exactly which extractor read the paper: all 94 came through GROBID, all 223 through the PDF’s own numbered list. doi_checker/doi.py rejects the prefix outright, and that rejection sits only in the path the XML extractor uses. So the same bibliography is judged differently depending on its layout. This is reported as found, not worked around: it is visible in RqInvalidAcmDl (164 of the 409 invalid DOIs) and it is what the RQ4 declined label counts.
doi-checker --version prints 0.1.0. Expected: the 1.0 release tag was cut without bumping doi_checker/version.py. The build that produced these results is identified by its commit, 5b40c64048b8, which is recorded in tools/doi-checker/COMMIT and in experiments/results-paper/tools.json and is what the tag points at.
hallucinator-cli reports missing shared libraries. Run ./setup_vm.sh, which installs them from vendor/deb/. ./setup_vm.sh --verify re-checks without changing anything.
A run sits at starting GROBID for minutes. Expected: the container loads its models before answering anything, measured at 2 min 21 s for one instance. GROBID_REPLICAS multiplies that and the memory, which is why it defaults to 1. NO_GROBID=1 skips it entirely and gives the same results, because the bundled extraction caches are complete. Interrupting during the wait is safe — the run says so and exits having done nothing.
A run is interrupted. Nothing is lost that had finished. The runner kills what was in flight, keeps every completed (tool, paper) pair, rewrites the CSVs from what is on disk, and prints how many it kept; the next run reports N (tool, paper) pair(s) already done; skipping and continues. --force redoes finished pairs instead.
A smoke test takes far longer than 3 minutes. If it is the first one on a cold machine, that is expected and happens once: the caches and mirrors are being read off physical disk through the shared folder. It has taken 37 minutes that way, against ~3.5 minutes when the host’s page cache is still warm from unpacking. If a later run is still slow, that is a different problem — see the next entry.
Everything is slow and the log is full of failures. Most likely a cache that is not being read. Run .venv/bin/python packaging/rekey_caches.py --check: it prints how much of the corpus each cache covers, and all four lines should read 100%. If refchecker’s reads 0%, the caches were not re-keyed after the artifact moved — re-run .venv/bin/python packaging/rekey_caches.py.
Two of those lines are about timestamps rather than about files being present, because doi-checker reuses a cached extraction only while the cache file is newer than the PDF, and a reproducible archive unpacks both with the same fixed timestamp — equal, which is not newer. install.sh fixes that by touching them; the --check report counts an entry that is not newer as not covered, so a package that was unpacked but never installed shows doi-checker pdftohtml XML 0 / 2500 rather than a reassuring 100%. Left unfixed it is quiet rather than loud: the tool simply re-runs pdftohtml, and the extraction then comes from whatever poppler is installed instead of from the shipped XML.
make check-corpus reports 1388 papers “not bundled”. That is the expected state of an unpacked package, not a fault: those papers’ licences do not permit redistribution, so they were never included. What matters is the “missing” column, which should be 0. See The corpus and its licences, and make corpus if you want the complete corpus.
A paper list names a paper that is not there. The runner refuses to start rather than silently running fewer papers than the list says. make check-lists checks both lists; make subset regenerates the subset. If the corpus itself is missing, see The corpus and its licences.
doi-checker reports a slightly different number of invalid DOIs than the paper. Expected, within a handful. An unregistered DOI is never cached (see the one difference), so every online run re-queries all 43 741 of them and a transient registry error leaves that one reference inconclusive rather than invalid. Three runs of 1.0 over the corpus gave 401, 403 and 452 invalid DOIs. Treat RqInvalidCount as a figure with a few units of run-to-run noise; the other two flag types come from cached records and are exactly reproducible.
A tool times out, or errors, on some papers. Recorded rather than swallowed: status is timeout or error for that one (tool, paper) pair and the run carries on, because “cannot finish this paper” is itself a datum. Over the paper’s own 3888-paper run that was 6 timeouts and 3 errors for hallucinator, 4 timeouts and 1 error for refchecker, and none at all for doi-checker — so a handful is normal and a flood is not.
make data-analysis warns about labels, environment or formats. Those three inputs are not produced by a run. The Makefile falls back to the paper’s copies under data-analysis/output-paper/, so a plain make data-analysis or make compare-paper finds them. The warning appears only where something bypasses the Makefile — such as the analysis step inside make smoketest, where RQ4 is not being checked anyway.
No Crossref --mailto given. Affects rate limits only, and only when running online. Copy .env.example to .env and set CROSSREF_MAILTO to move into Crossref’s polite pool. Every key in that file is optional; none is needed for anything in this README.The Teop language documentation and the content of this volume
Design, construction and validation of an open-source gold nanoparticle concentration tracking device
ТРАНСКУЛЬТУРНО-СОПОСТАВИТЕЛЬНОЕ ИССЛЕДОВАНИЕ ИСКУССТВЕННОГО ИНТЕЛЛЕКТА И ПЕРЕВОДЧЕСКОЙ ПРАКТИКИ В АНГЛО-УЗБЕКСКОМ ХУДОЖЕСТВЕННОМ ПЕРЕВОДЕ
Analysis code for "A Multi-Omics Brain-Heart Interactome Maps Signaling Networks Associated with Post-Stroke Cardiac Dysfunction"
Co-localized scanning probe microscopy-Raman scattering studies of hybrid plasmonic substrates for SERS
Replication package for "Replication package: simulator, per-topology records and analysis code for learned route scoring in wireless sensor networks"
The Teop people and their language
Development of a COSMO-SAC Parametrization with Advanced QM Method TZVPD-FINE
MANAGING TRUST, COMMUNICATION, AND PERFORMANCE IN CROSS-CULTURAL VIRTUAL TEAMS THE ROLE OF CULTURAL INTELLIGENCE A Conceptual Model and Exploratory Workplace Survey
Reproduction Package for ATIQSER 2026 Proceedings "doi-checker: A Tool for Validating References of Research Papers"
On Losses, Pauses, Jumps and the Wideband E-Model – IEEE Xplore Document
There is an increasing interest in upgrading the EModel, a parametric tool for speech quality estimation, to the wideband and super-wideband contexts. The
NUAV – a testbed for developing autonomous Unmanned Aerial Vehicles – IEEE Xplore Document
Contemporary models of Unmanned Aerial Vehicles (UAVs) are largely developed using simulators. In a typical scheme, a flight simulator is dovetailed with a
NUAV – a testbed for developing autonomous Unmanned Aerial Vehicles
Simulators as Drivers of Cutting Edge Research – IEEE Xplore Document
Undertaking engineering research can be compounding for beginning graduate students and thwarting even for seasoned researchers. With a wealth of academic
Simulators as Drivers of Cutting Edge Research
Evolutionary speech quality estimation in VoIP
A Methodology for Deriving VoIP Equipment Impairment Factors for a Mixed NB/WB Context
Real-Time, Non-intrusive Speech Quality Estimation: A Signal-Based Mod
Real-Time, Non-intrusive Evaluation of VoIP
VoIP speech quality estimation in a mixed context with genetic programming
An Evolutionary Approach to Speech Quality Estimation
Real-Time Non-Intrusive VoIP Evaluation Using Second Generation Network Processor
Non-intrusive quality evaluation of VoIP using genetic programming
