Blog

IOI Fund for Network Adoption: LA Referencia’s six-month progress update

Network adoption with a strong technical foundation.

August 31, 2026 · 8 min read

This article is also available in Spanish.

In late 2025, Invest in Open Infrastructure launched the IOI Fund for Network Adoption, a new kind of investment designed to accelerate the global adoption of open infrastructure for research and data sharing. Rather than funding individual tools or single institutions, the Fund invests in networks: consortia of institutions and researchers with established relationships, active member governance, and a demonstrated commitment to open science. The logic is that strengthening the shared infrastructure that a network already runs can multiply impact across every member institution at once.

The Fund grew out of direct engagement with research communities to identify the open infrastructure projects most in need of sustained support, and it pools contributions from a group of funders all interested in strengthening the global open research ecosystem. Each selected network receives up to $1.5 million over two to three years, paired with ongoing implementation support from IOI staff on governance, community engagement, and business modeling. IOI’s goal with the Fund is not only to enable grantees to do projects that can have a big impact on their networks, but also to support the resilience of these networks long after the grant ends.

LA Referencia was named one of the Fund's two inaugural grantees, selected from more than 100 applications spanning 22 countries. Together, the Fund's first two grantees are positioned to benefit more than 1,300 research institutions across 36 countries in Latin America and Africa. This post provides an overview of what LA Referencia has done with the first six months of that investment. Look for an update post about our other inaugural grantee, UbuntuNet Alliance, in the coming months. 

Logo of LA Referencia

An introduction to LA Referencia’s project

LA Referencia is a regional cooperation network that aggregates and provides open access to the scholarly and scientific output of national repositories across Latin America and Spain, built on a model of federated governance and publicly owned, non-commercial infrastructure. With IOI Fund support, LA Referencia is running a 2026–2028 project to expand that infrastructure's regional reach and technical depth with four major goals: bringing more countries and repositories into the network, modernizing how researchers discover content in multiple languages, building a sovereign and verifiable system of persistent identifiers, and standing up a shared regional repository and training program for research data. The project's full progress is documented in detail on LA Referencia's project reporting site, and the account below draws directly on the team's own executive report for January–July 2026

The project is organized into three complementary components that connect content, discovery technology, identifier infrastructure, and human capacity:

  • Component A — Expanding LA Referencia’s Aggregation Infrastructure: growing regional coverage, improving and normalizing metadata, and introducing multilingual semantic search to LA Referencia's discovery layer.
  • Component B — Decentralized ARK Identifiers and Metadata Preservation (dARK): building dARK 2.0, a modular, blockchain-backed infrastructure for assigning, publishing, and resolving persistent identifiers.
  • Component C — Regional Orphan Data Repository & Dataverse Support: standing up a Dataverse-based regional repository for research data alongside a federated training and curation program.

In the rest of this post, we dive deeper into the progress made on each of these components and share explorable outputs of LA Referencia’s work to date. 

Component A: Expanding LA Referencia’s Aggregation Infrastructure

Discovery depends on good metadata for the collections, so Component A pairs the expansion of LA Referencia's regional coverage with metadata quality work and a new multilingual search capability that enables discovery regardless of the language of the search or the content.

Progress to date

Over the reporting period, LA Referencia enacted a regional methodology for identifying candidate institutions and repositories for harvesting to bolster their collection. This work involved verifying OAI-PMH services, reviewing metadata, updating COAR's Interoperability and Registry Database (IRD), and cross-checking institutional information against ROR. Curation work was completed for member countries Ecuador, Uruguay, and Costa Rica, and non-member countries Colombia, Paraguay, Venezuela, Bolivia, and Cuba; records for another countries are in progress. 

In total, the team updated records for 362 repositories: 187 edited, 77 newly added, and 98 archived. Venezuela was the country with the greatest progress, with two repositories successfully harvested as a pilot for incorporating content from countries that do not yet operate as full nodes in LA Referencia. 

Related to this geographic expansion, LA Referencia is building a solution for a common problem they encountered in their community: many researchers search in one language while the work they need is described in another. A record from a Spanish-language repository might never surface for someone searching in English, or vice versa, because traditional keyword search can only match the exact words on the page. Component A set out to close that gap with multilingual semantic search, which finds results by matching meaning instead of exact wording. To figure out the best way to do this, the team ran a rigorous, reproducible evaluation, comparing roughly thirty different AI models across two test collections of 30,000 to 50,000 documents, with thousands of test searches in as many as 10 languages. This work resulted in promising outcomes: in the pilot, semantic and blended search surfaced roughly twice as many relevant results in the top 10 (up from an average of 4.5 to 9.1) compared with keyword search alone, with the biggest gains for languages that keyword matching tends to handle poorly, like Chinese, Japanese, and Arabic. The team evaluated different models and chose one that balanced high-quality results with efficiency and speed to run on millions of records. 

That evaluation has now been built into a beta working system on LA Referencia's search platform. When someone runs a search, they can choose the traditional keyword match, a "meaning-based" match that works across languages, or a blended mode that combines both. Reindexing becomes substantially faster when previously generated embeddings are reused. Rebuilding the full index for a test set of about 28,000 records took roughly 5 hours the first time, but only about 10 minutes once the meaning-based data was already stored, making regular updates practical as the collection keeps growing. Alongside this, the team added a open WAF that screens out automated bots and scrapers, plus a way to apply that same protection across multiple LA Referencia sites at once.

You can also dive into the technical details by reading LA Referencia’s full Component A report.

What's next

LA Referencia continues to review and add repositories from additional member countries and new countries beyond their network. The team plans to expand its multilingual semantic search evaluation across the full regional collection and consolidate the semantic search beta ahead of broader rollout.

Component B: Decentralized ARK Identifiers (dARK)  and Metadata Preservation

Component B involves building dARK 2.0, the next generation of LA Referencia's infrastructure for assigning, publishing, and resolving persistent ARK identifiers.  This component accounts for regional, sovereign, and sustainable infrastructure built specifically for the region served by LA Referencia.

Progress to date

Persistent identifiers are what keep a piece of research findable and citable for the long term by using a permanent link that still leads to the right paper, dataset, or report years later, even if the file moves or the system hosting it changes. Today, most of that infrastructure is run by a handful of external providers. dARK 2.0 is LA Referencia's vision for an infrastructure the region can run, govern, and trust for itself without relying on one single organization as a locus of control.

To do that, the team split dARK 2.0 into separate components that can each be checked, updated, or scaled independently. A shared, tamper-evident ledger (a type of blockchain) maintains the authoritative record of which identifiers exist and what they point to, while the descriptive metadata for each record is stored separately, allowing anyone to verify it hasn't been altered. That split keeps the "proof" layer simple and trustworthy, while the descriptive data can keep growing without needing to touch the "proof" layer. The team also built the full pipeline this system requires: reserving new identifiers, preparing and storing their metadata, publishing the reference to the ledger, and allowing anyone to look up an identifier and resolve it to the correct record.

The team stress-tested this workflow against millions of records and ran hundreds of automated checks to confirm it holds up, and built administrative tools so staff can monitor identifiers, activity, and lookups. The dARK pipeline is now ready to be deployed and connected to LA Referencia's regional harvesting system, so the harvesting workflow is being integrated with the dARK pipeline to support identifier assignment as records are processed, even across large, constantly changing collections. 

Dive into the details in the full Component B report.

What’s next

The next step is spreading the dARK infrastructure across institutions rather than running it from one place. A matching setup is being prepared at RNP in Brazil, and partner institution IBICT is preparing to migrate its identifiers over from the original 1.0 version of dARK. With multiple institutions using the system, LA Referencia plans to validate cross-site operation and formalize the security and governance model for this federated regional infrastructure.

Component C: Regional Orphan Data Repository & Dataverse Support

Component C pairs a new Dataverse-based regional repository with a federated training and curation program, aimed at giving research data lacking institutional infrastructure a home while spreading the skills needed to run data services sustainably across the region.

Progress to date

The team began with a regional diagnostic that identified roughly 96 registered data repositories and around 30 Dataverse installations across Latin America, concentrated in Brazil, Argentina, Mexico, and Colombia, drawing on questionnaires and interviews covering governance, infrastructure, curation, preservation, support, and interoperability. That diagnostic confirmed the region already has significant experience and expertise, which has shaped a peer-collaboration strategy treating the repository as a socio-technical service: software and infrastructure paired with policy, multidisciplinary teams, researcher support, and ongoing training.

The headline technical result is a live, working test version of the repository (a "Dataverse Alpha") running on LA Referencia's own infrastructure. The repository works in English, Spanish, and Portuguese, includes fields to tag a dataset's country of origin and which of the UN's 17 Sustainable Development Goals it relates to, and uses open licenses (CC0 and CC BY) so the data can be freely reused. Sample datasets are already loaded so the team can test uploading, describing, searching, and accessing files before real institutional deposits start coming in.

Alongside the technical build, the team is developing a 16-hour asynchronous MOOC, Introduction to Data Management and Curation in Dataverse, covering foundations and publication, metadata and documentation, technical verification and formats, and ethics, licensing, and approval. This is one of three planned training products; the others include hands-on curation training and direct support for pilot institutions. The course will be built and installed on a production Moodle environment with materials in Spanish and Portuguese.

We encourage readers to explore the full technical details in the full Component C report.

What’s next

The team will finish course content, validate repository workflows end-to-end, select the first pilot institutions, and turn the lessons from the Alpha into a stable model for the regional service.

LA Referencia’s Project: Supporting the Goals of the IOI Fund for Network Adoption

Six months in, LA Referencia has moved from design and diagnostic work to functioning systems: a measurably better multilingual discovery experience, a working architecture for verifiable, sovereign persistent identifiers, and a live regional data repository paired with a training pipeline built to outlast any single grant. None of this was built for one institution; every improvement contributes to shared, federated infrastructure that LA Referencia's member repositories across Latin America already rely on. 

That is the premise of the IOI Fund for Network Adoption: that the fastest, most durable way to grow open infrastructure adoption is to invest in networks that already have the governance, trust, and reach to put new capacity to work across many institutions at once. LA Referencia's first six months show that thesis in motion. The technical foundation that they are building will enable the project's next phase, expanding to new countries and languages, validating dARK across institutions, and bringing its first pilot institutions onto the regional Dataverse.


Read more about the project on LA Referencia's project reporting site, including the full executive report this post is based on, or learn more about IOI's work with LA Referencia and the IOI Fund for Network Adoption

You’ve successfully subscribed to Invest in Open Infrastructure
Welcome back! You’ve successfully signed in.
Great! You’ve successfully signed up.
Your link has expired
Success! Check your email for magic link to sign-in.
Please enter at least 3 characters 0 Results for your search