# ECAI Indexing: Turn Source Collections into Managed Work

A practical guide to ECAI's durable indexing queue, JSONL and IPFS inputs, disk indexes, ingestion journal and separate private retrieval path.

Canonical HTML: <https://damagebdd.com/articles/ecai_indexing_workflows.html>



<a id="an-index-needs-an-operating-workflow"></a>

## An index needs an operating workflow

Building a search index once is straightforward compared with keeping one
useful. Sources change, requests are retried, a worker stops halfway through
an import, and someone needs to explain whether a completed job produced
searchable data or merely recorded receipt of it.

ECAI addresses these operational questions with a durable indexing queue and
adapters for several source formats. A job has an owner, a source description,
progress events and an output artifact. That gives teams a common place to
observe work even when the underlying storage paths differ.

For an adopter, the useful feature is control: choose a bounded collection,
follow its progress, recover interrupted work and inspect the output before
making it part of a search service.


<a id="choose-the-output-as-carefully-as-the-input"></a>

## Choose the output as carefully as the input

The reviewed implementations do not all produce the same kind of index.

| Input and path | Result | Best starting use |
| --- | --- | --- |
| Wikipedia JSONL or Yelp NDJSON | Records in shared live search | Existing structured exports |
| Wikimedia visibility pipeline | Isolated live-search snapshot and corpus material | A selected public reference corpus |
| IPFS CID or manifest, `searchable_disk` | Published posting segments and document store | A public collection retained on disk |
| IPFS CID or manifest, `ledger_only` | Durable ingestion records | Recording accepted input for later processing |
| Private-index API | Encrypted segments under corpus permissions | A separately controlled internal collection |

The private API is a separate path, not another public queue mode. Public
indexing guards reject records marked private. Read the
[private knowledge guide](ecai_private_knowledge.md) before bringing confidential material into a pilot.

A ledger acknowledgement is also different from a searchable result. The
IPFS adapter's journal path records durable ingestion, while artifact
finalisation rejects `ledger_only` as a searchable index. A downstream
consumer still has to turn those records into queryable material.


<a id="give-long-running-jobs-an-owner-and-a-lifecycle"></a>

## Give long-running jobs an owner and a lifecycle

The [queue service](https://github.com/DamageBDD/DamageBDD/blob/a993c0421b38de475b103dc79f31da1e091cd97a/apps/ecai/src/ecai_index_jobs_srv.erl) stores jobs and events through a disk-backed store. It supports
priority ordering, configurable concurrency and queue limits, owner-scoped
idempotency, and pause, resume, cancel and retry operations. Repeating a
submission with the same idempotency key can be distinguished from a conflicting
request rather than silently creating another job.

Workers report progress to the queue. On service recovery, interrupted work
is returned to an appropriate recoverable state, and reports from superseded
workers are fenced out. The exact amount of work repeated depends on the
adapter and its checkpoints; durable queue state is not a universal promise
of exactly-once ingestion.

The [HTTP interface](https://github.com/DamageBDD/DamageBDD/blob/a993c0421b38de475b103dc79f31da1e091cd97a/apps/ecai/src/ecai_index_jobs_http.erl) exposes job submission and listing at `/ecai/index-jobs`,
per-job controls, artifact retrieval and a progress-event feed. It uses the
existing authenticated account path and scopes job access to the owner.
Wikimedia presets let the server retain source and storage policy while the
client chooses a named configuration.

This is the operational index queue. The similarly named `ecai_jobs_srv`
module is a separate chunk-job prototype discussed in the
[index NFT article](ecai_index_nfts.md). Its state and payment labels should not be confused with
the durable queue's guarantees.


<a id="keep-the-connection-between-a-chunk-and-its-source"></a>

## Keep the connection between a chunk and its source

For IPFS input, an operator can supply one content identifier or a manifest
describing multiple documents. The adapter fetches the material, reports
failures with the document's position and identifier, and sends accepted
records to the selected ingestion path.

The [shared chunker](https://github.com/DamageBDD/DamageBDD/blob/a993c0421b38de475b103dc79f31da1e091cd97a/apps/ecai/src/ecai_chunker.erl) creates versioned UTF-8 text windows with overlap and byte
offsets. A separate NDJSON mode keeps line records as chunks. Ingestion
records retain source, content and chunk identities, giving later processing
a way to relate a search excerpt to the material that produced it.

This is useful for manuals, public incident reports and structured directories:
an answer can point back to a particular source version instead of an
unidentified excerpt. Chunk size and overlap still affect relevance, storage
and duplicate results, so they belong in evaluation rather than being treated
as universally optimal settings.

For local structured files, source descriptors include their order, length
and digest. Finalisation rechecks the files and rejects changed input. Local
absolute paths are excluded from the portable source identity.


<a id="disk-storage-has-its-own-publication-boundary"></a>

## Disk storage has its own publication boundary

The [disk indexer](https://github.com/DamageBDD/DamageBDD/blob/a993c0421b38de475b103dc79f31da1e091cd97a/apps/ecai/src/ecai_disk_indexer.erl) writes bounded batches and synchronises its document store
before making a segment visible through the index manifest. A failed later
batch does not imply that earlier published segments disappear. Similarly,
a multi-document IPFS import is not one atomic transaction covering every
document.

The [disk search module](https://github.com/DamageBDD/DamageBDD/blob/a993c0421b38de475b103dc79f31da1e091cd97a/apps/ecai/src/ecai_disk_search.erl) retrieves and merges a term's document identifiers across
published segments, with a cache for frequently used terms. This is a useful
search building block. It does not by itself expose the same ranked,
field-aware result interface as the in-memory search module.

Artifact contents also differ. Wikimedia jobs package a complete search
snapshot; legacy Wikipedia JSONL and Yelp jobs export term headers from the
shared live context. Those headers alone are not a portable, dataset-exclusive
copy of all records imported by that job.


<a id="a-practical-route-to-adoption"></a>

## A practical route to adoption

Start with public records whose expected results are easy to recognise.
Submit the same job twice, inspect idempotency behaviour, pause and resume it,
restart the service, and verify the output that a separate consumer can load.
Measure the disk and memory required by the chosen mode.

The repository's [queue tests](https://github.com/DamageBDD/DamageBDD/blob/a993c0421b38de475b103dc79f31da1e091cd97a/apps/ecai/test/ecai_index_jobs_srv_tests.erl) and
[artifact tests](https://github.com/DamageBDD/DamageBDD/blob/a993c0421b38de475b103dc79f31da1e091cd97a/apps/ecai/test/ecai_index_artifact_tests.erl) provide starting cases to inspect and run in a matching
environment. This article reports source review, not a completed runtime pilot.

Once that workflow is measured, a team could reuse it for scheduled public
document releases, independently built reference collections or an ingestion
service shared by several applications. Each new source adapter would need
its own identity, recovery and publication contract. The common queue makes
those responsibilities visible.

