Most of our work has resulted in scholarly publications. On this page you can review our publications to get an idea about our work.
India's Software Stories: Preserving code as heritage with Software Heritage, Commons & Wikidata
July, 2026 • Presentation
Tosoni, Francesco, Surampudi, Pavan Santhosh
Slides of the lightning talk India's Software Stories, presented at Wikimania 2026 (Paris, France) on 23 July 2026, in the Lightning Talk Showcase I.
Software is fragile knowledge: as the UNESCO Paris…
Slides of the lightning talk India's Software Stories, presented at Wikimania 2026 (Paris, France) on 23 July 2026, in the Lightning Talk Showcase I.
Software is fragile knowledge: as the UNESCO Paris Call on Software Source Code (2019) puts it, source code is part of our shared documentary heritage. The talk shows how three open infrastructures give the world's code stories a permanent home: Software Heritage preserves every publicly available line of source code; Wikidata links it through citable persistent identifiers (SWHIDs, property P6138); Wikimedia Commons holds the human story: people, places, screenshots and video. Indian examples include Lekhini (typing Telugu on any English keyboard), UrduScript, the Slug GPU rendering algorithm and Ezhil, a programming language whose keywords are Tamil words (Q12975625). The talk invites Wikidata editors, GLAM institutions, digital-preservation and FOSS practitioners, and language and Global South communities to document their region's technological and linguistic history through the Software Stories project.
Links
Video recording (YouTube): youtube.com/live/saeV6bZ7cIk (from 6:58:25)
Wikimedia Commons: original upload of the slides
Software Heritage Stories: stories.softwareheritage.org
Third-party images are credited on each slide under their own licences.
MediaWiki Code2Code Search: Semantic Search to Find Code by Under-the-Surface Similarity
May, 2026 • Presentation
Tosoni, Francesco
Slides of the talk presenting MediaWiki Code2Code Search, given at the Wikimedia Hackathon 2026 in Milan (Italy) on 2 May 2026, 11:15–12:00 (programme; Phabricator task T425057).
Unlike classical keyw…
Slides of the talk presenting MediaWiki Code2Code Search, given at the Wikimedia Hackathon 2026 in Milan (Italy) on 2 May 2026, 11:15–12:00 (programme; Phabricator task T425057).
Unlike classical keyword-based code search, Code2Code Search takes a code snippet as the query and retrieves semantically similar implementations across 2,400+ MediaWiki repositories, regardless of variable names, coding style or exact syntax, which is useful for maintenance and vulnerability tracing. Code is mapped to embeddings and retrieved with FAISS; the index covers 1.1M+ snippets from 83k+ source files in ten programming languages, built offline on Toolforge. All indexed repositories are archived in Software Heritage, so every result can be cited through a persistent, intrinsic SWHID. The multilingual interface supports 14 Indic languages alongside Italian and French. The talk closes with the roadmap: dynamic indexing, a transition to the Wikimedia Codex design system, broader language support and community feedback.
Links
Tool (Toolforge): code2codesearch.toolforge.org
Source code: github.com/ftosoni/mediawiki-code2code-search (software record: 10.5281/zenodo.20586244)
Toolhub: toolforge-code2codesearch
Wikidata: Q139251277
Wikimedia Commons: Category:MediaWiki Code2Code Search
Diff post: Introducing Mediawiki Code2Code Search (14 April 2026)
The slides were originally uploaded to Wikimedia Commons: File:Mediawiki-code2codesearch-wmhackathon2026-slides.pdf. Third-party images are credited on each slide under their own licences.
Supplementary Data for How robust is short-run GDP–CO2 coupling in Thailand? A small-sample time-series reassessment, 2000–2024
October, 2026 • Dataset
podong, chattanong
Supplementary dataset supporting the study How robust is short-run GDP–CO2 coupling in Thailand? A small-sample time-series reassessment, 2000–2024. The dataset contains annual data used f…
Supplementary dataset supporting the study How robust is short-run GDP–CO2 coupling in Thailand? A small-sample time-series reassessment, 2000–2024. The dataset contains annual data used for the analysis of the short-run relationship between economic activity and carbon dioxide emissions in Thailand over 2000–2024. Variables used in the study include CO2 emissions, population, GDP per capita, derived CO2 emissions per capita, and variables used in the robustness analyses. The dataset is provided to support transparency, reproducibility, and independent verification of the reported results.
Thailand CO2 emissions Economic growth GDP per capita Carbon emissions Time-series analysis Robustness analysis Small-sample inference Climate policy
PRENDIZAGEM ATIVA DA PEDOLOGIA NA EDUCAÇÃO BÁSICA : Análise da Oficina "Solos em Ação: Experiência para Sentir e Aprender"
October, 2026 • Conference paper
Rodrigues Barros, Lucas Guilherme, Souza, Hillary da Silva, Bento, Fabriele Valeska dos Santos
RESUMO: O ensino de Geografia enfrenta desafios significativos, especialmente no que se refere àarticulação entre teoria e prática no ensino da pedologia em sala de aula. A…
RESUMO: O ensino de Geografia enfrenta desafios significativos, especialmente no que se refere àarticulação entre teoria e prática no ensino da pedologia em sala de aula. Abordagensexcessivamente teóricas dificultam a compreensão do solo enquanto elemento fundamental dasrelações entre sociedade e natureza. Diante desse contexto, o presente artigo tem como objetivoapresentar e analisar uma experiência pedagógica desenvolvida com alunos do 6º ano do EnsinoFundamental II, por meio de uma oficina de Educação dos Solos, fundamentada emmetodologias práticas, experimentais, sensoriais e táteis. Sob essa Ótica, pesquisa possuiabordagem qualitativa e foi realizada por meio de uma oficina pedagógica de caráter teóricoprático, com ênfase nas atividades práticas. As ações envolveram a exposição e a observaçãode diferentes tipos de solo, como solo argiloso, solo húmico, Gleissolo e Neossolo Fluvial, alémda apresentação de solo indiferenciado coletado em cemitério, com finalidade comparativaentre solos com horizontes definidos e não definidos. Os estudantes também participaram dacoleta de solos no ambiente escolar e de atividades no “Laboratório do Agricultor”, onderealizaram experimentos de decantação, possibilitando a identificação de materiais orgânicos,inorgânicos e partículas de diferentes densidades. Logo,os resultados evidenciaram maiorengajamento dos estudantes e avanços significativos na compreensão das características dossolos, especialmente no que se refere à fertilidade, textura e composição. Conclui-se que o usode metodologias ativas no ensino de Geografia contribui para uma aprendizagem maissignificativa, fortalecendo o pensamento geográfico e a consciência ambiental dos alunos,contribuindo para práticas pedagógicas inovadoras no contexto escolar contemporâneobrasileiro.
Ensino de GeografiaPedologiaMetodologias AtivasEducação dos SolosEnsino Fundamental II
A DEGRADAÇÃO DOS RIOS E NASCENTES DE ARAÇOIABA-PE: OS ESTUDANTES DA EDUCAÇÃO BÁSICA COMO PESQUISADORES DA HISTÓRIA LOCAL
October, 2026 • Conference paper
Silva, Gabriel Arcanjo de Carvalho, Mendes de Andrade, Alex, Rêgo, Ananda do Nascimento
RESUMO: É fundamental para estudantes do ensino básico a compreensão territorial, social eambiental de sua localidade, assim como, é dever dos educadores o ensino desses si…
RESUMO: É fundamental para estudantes do ensino básico a compreensão territorial, social eambiental de sua localidade, assim como, é dever dos educadores o ensino desses significados.Em face disso, o presente artigo nasce de uma experiência pedagógica, vivenciada na EscolaMunicipal Dom Pedro II, na cidade de Araçoiaba-PE, sendo o projeto um recorte de umaatividade desenvolvida pelos estudantes da rede para a Semana Municipal de Ciência eTecnologia de Araçoiaba-PE 2025, cujo tema central foi: “Planeta Água: a cultura oceânicapara enfrentar as mudanças climáticas no meu território”. Nesse sentido, um grupo de estudantesdos anos finais do colégio, orientados pelo professor de História e Cidadania, com o objetivode protagonizar o papel de pesquisadores da comunidade onde vivem, fizeram entrevistas comseus avós, pais e demais habitantes, buscando relatos orais sobre os rios, nascentes e matasciliares do território. A metodologia partiu de um questionário semiestruturado com perguntasa respeito das relações socioambientais da comunidade para com as águas, da condição em quese encontram esses ambientes e se são preservados ou não. Os resultados obtidos tiveramconsonância com o que o Grupo de Estudos e Pesquisa sobre Histórias e Geografias deAraçoiaba-PE (GPHG) define como um extenso problema para o território, dos quais sedestacam a precariedade administrativa na conservação ambiental e a presença maciça doagronegócio monocultor de cana-de-açúcar. A consciência dessas adversidades na comunidadeainda se encontra em fase embrionária, sendo esse um desafio educacional, reforçando assim opapel da escola como espaço tático para a construção de saberes socioambientais coletivos.
Ensino de HistóriaEducação AmbientalRios e Nascentes
Introducción: La aplasia cutis congénita (ACC) es un trastorno poco frecuente que se caracteriza por la ausencia localizada de piel al nacimiento, que compromete principalmente el cuero …
Introducción: La aplasia cutis congénita (ACC) es un trastorno poco frecuente que se caracteriza por la ausencia localizada de piel al nacimiento, que compromete principalmente el cuero cabelludo, con lesiones que pueden ser desde superficiales hasta extensas asociadas a compromiso óseo o sistémico. Metodología: Se presenta un caso clínico de un recién nacido de término (RNT) hijo de madre con diabetes gestacional (DG), nacido por cesárea de urgencia que presenta una lesión extensa de ACC del cuero cabelludo. La evaluación clínica consideró la extensión, localización y profundidad de la lesión, así como la presencia de compromiso neurológico, óseo, vascular o sistémico. Para orientar la conducta terapéutica se realizó una revisión narrativa de la literatura disponible sobre ACC, enfocada en criterios de evaluación, factores asociados a complicaciones y alternativas de manejo conservador y quirúrgico. Resultados: Se realizó evaluación multidisciplinaria sin evidenciar compromiso sistémico asociado. El paciente evolucionó favorablemente con manejo conservador. Discusión: La extensión de la lesión y la cercanía con estructuras vasculares críticas representaron un riesgo clínico significativo; sin embargo, la ausencia de compromiso óseo identificable mediante radiografía, junto con la estabilidad clínica y la ausencia de complicaciones hemorrágicas o infecciosas, permitió mantener un manejo conservador bajo vigilancia estrecha. Conclusión: Este caso destaca la importancia de un enfoque integral en la evaluación de la ACC extensa, así como el rol del manejo conservador en pacientes seleccionados sin complicaciones mayores.
Reproducibility package: Credential expansion and adult foundational skills in Chile (SIES 2007-2023 and PIAAC 2014/15-2022/23)
October, 2026 • Dataset
Suazo Galdames, Ivan, Chaple Gil, Alain Manuel
What this is. The complete data, code, results and audit trail behind the studyCredential expansion and adult foundational skills in Chile: system trajectory 2007–2023 and change in adultliterac…
What this is. The complete data, code, results and audit trail behind the studyCredential expansion and adult foundational skills in Chile: system trajectory 2007–2023 and change in adultliteracy and numeracy across PIAAC cycles. The study asks how far the expansion and changing composition ofChilean higher education credentials corresponds with change in measured adult literacy and numeracy, keeping threetime windows explicitly separate: the trajectory of the system (2007–2023), the interval between the twoassessments (2014/15 to 2022/23), and the contemporaneous correspondence within that interval. The design isecological and supports no causal inference.
Data sources. Three bodies of evidence. First, the complete administrative record of the ChileanServicio de Información de Educación Superior (SIES): annual credential counts and system-wide enrolmentfor every year of the period. Second, the public use files of both comparable cycles of the OECD Programme for theInternational Assessment of Adult Competencies for Chile, reanalysed at the individual level for the protocol'spre-specified population of adults aged 25 to 65. Third, published results for 27 participating countries andeconomies, reanalysed on an annualised basis because elapsed intervals between cycles range from 6.0 to 11.0 years.
Contents.
00_protocol/ the dated pre-analysis protocol (2026-09-11), frozen but not publicly registered.01_source_manifests/ origin, size, SHA-256 digest, access date and analytic role of all 25 source files.02_derived_data/ 36 CSV files: harmonised, analytic and revision datasets, including the Chileanmicrodata estimates under both populations and both exclusion variants, the enumerated 39-contrast multiplicityfamily under two standard-error conventions, the annualised international meta-analysis and meta-regression, and theSIES-to-ISCED credential crosswalk with flow shares by class and year.03_code/ the analysis scripts (Python and R), with the superseded versions retained under asuperseded_ prefix.04_results/ the nine published tables, the eight published figures at 300 dpi in PNG and PDF, and thesource data for every figure.05_manuscript/ and 06_supplement/ the revised manuscript, supplementary material and title page.07_audit/ the audit trail: 36 automated checks tying every headline number in the manuscript to itssource table, a point-by-point response to a thirty-point methodological review, per-module methods memos, variabledictionaries and the decision log.
The standard-error rule that governs every inference here. A between-cycle difference computed fromthe public use files carries sampling and imputation variance only. It omits the error of linking the two cycles'reporting scales, which the publisher estimates separately and folds into its published trend standard errors. Acrossthe twelve statistics at ages 16 to 65 where both exist, the microdata standard error is 40 to 93 per cent of thepublished one, with a median near 85 per cent. The published standard error is therefore primary wherever one exists;microdata standard errors are primary only where no published counterpart exists, and those contrasts areanti-conservative. Every contrast at ages 25 to 65 is additionally reported under a linking-error-adjustedapproximation. Anyone reusing the microdata standard errors directly will overstate significance.
Replication notes. Cycle 1 uses jackknife-2 variance, Cycle 2 uses Fay balanced repeatedreplication with a factor of 0.3. The 80 replicate columns are padded with duplicates of SPFWT0: the Fay denominatormust use 28 effective replicates in Cycle 2 and 17 in Cycle 1, not 80, or the sampling variance is understated and thepublished standard errors will not reproduce. Computed at ages 16 to 65 with no cross-cycle exclusions, the pipelinereproduces all twelve published Cycle 1 statistics to within 4e-08 and the Cycle 2 statistics to within 0.013 points.A fixed seed (20260911) governs the bootstrap.
What is not here. The two PIAAC public use files are not redistributed. Their distribution URLs,access dates and SHA-256 digests are recorded in source_manifest.csv so that an identical copy can beobtained from the publisher and verified. No record-level link exists between any credential record and any assessedrespondent, and none is asserted anywhere in this package. No renv lockfile is supplied, because writingone by hand would assert a dependency graph that was never resolved. The Module A extraction stage was not re-executedand is covered instead by the checksum manifest, external validation against official reports and the internal audit.
Known constraints. Read 07_audit/response_to_review.md and section S1 of thesupplement before reusing these outputs. The ISCED flow classification is bounded rather than exact, at category leveland not validable at person level; the assignment of diplomado and postítulo to the non-ISCEDclass is the load-bearing judgement. The decomposition is published for ages 16 to 65 only. The cohort analysis is notidentified: exposure and ageing remain collinear by design and the most exposed birth cohorts are structurally absent.Regional composition measures provision rather than residence, and distance-mode credentials rose from 0.91 to 16.63per cent of the flow between 2007 and 2023, so part of the measured metropolitan concentration is a registrationartefact.
Vocabulary Selection Strategies for Protein Language Modeling
March, 2026 • Thesis
Sargsyan, Maria, Klakow, Dietrich
Large Language Models have extended beyond natural language processing tasks such as
language generation, becoming transformative tools in advancing the field of structural
biology. However, tradition…
Large Language Models have extended beyond natural language processing tasks such as
language generation, becoming transformative tools in advancing the field of structural
biology. However, traditional large language models (LLMs) predominantly employ
subword tokenization techniques (e.g., Byte-Pair Encoding), followed by complex pre-
tokenization schemes. In contrast, most protein language models (PLMs) typically rely on
character-level tokenization, treating each amino acid individually. This design choice is
often a byproduct of the downstream tasks of interest, which are predominantly defined
at the residue level.
Our key findings reveal a divergence between the results of intrinsic and extrinsic
evaluations. Amino acid-level PLMs perform better in terms of normalized perplexity
on a test partition of the pretraining data. Among the intrinsic metrics, minimum
description length is the one that most closely aligns with downstream performance,
favouring a vocabulary size of around 4,000 tokens for proteins and larger vocabularies
for natural language. Meanwhile, an intermediate vocabulary size of 1,000 is optimal for
Deeploc’s downstream classification tasks. The results suggest that tokenization in PLMs
is task-dependent. We further demonstrate that BPE tokenization recovers biologically
meaningful sequence patterns such as nuclear localization signals and hydrophobic
repeats. Additionally, we examine the impact of pre-tokenisation on the pre-training of
LLMs and present evidence that the pre-tokeniser is an effective regularization tool, a
prior that is currently absent in PLMs.
There is an increasing interest in upgrading the EModel, a parametric tool for speech quality estimation, to the wideband and super-wideband contexts. The
Contemporary models of Unmanned Aerial Vehicles (UAVs) are largely developed using simulators. In a typical scheme, a flight simulator is dovetailed with a
Undertaking engineering research can be compounding for beginning graduate students and thwarting even for seasoned researchers. With a wealth of academic
To provide the best experiences, we use technologies like cookies to store and/or access device information. Consenting to these technologies will allow us to process data such as browsing behavior or unique IDs on this site. Not consenting or withdrawing consent, may adversely affect certain features and functions.
Functional
Always active
The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network.
Preferences
The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user.
Statistics
The technical storage or access that is used exclusively for statistical purposes.The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you.
Marketing
The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes.