Copal: a file service on one database
Almost everything I build ends up needing somewhere to put files. Antumbra keeps the original bytes of every document it learns from. Penpal is meant to archive the design exports it brokers. PolyConsoleA game platform I'm building, with apps on Android and an online social service behind them.'s players are going to upload things. Each time, the list of wants is the same. Keep the bytes, and don't keep them twice. Check what they actually are. Pull the text out so I can search it. Keep the old versions. Hand out a link that stops working. Tell me when it's done.
The usual way to get that list is to assemble it: an Object storeStorage that keeps files as whole objects under names, like Amazon S3, rather than as folders on one disk. for the bytes, PostgresPostgreSQL, a widely used open-source relational database. for the MetadataData about data. For a file, its name, type, size, owner and history, as opposed to its contents., a Job queueA waiting line of jobs that background workers pick up and process one by one. and WorkerA background process that picks up jobs and does them, apart from the requests people are waiting on. for the processing, a search engine for the text, and something append-only for the Audit trailA permanent record of who did what and when, kept so it can be checked later and cannot be quietly rewritten.. That's five systems to run, and five places for the truth about one file to disagree. The upload finished but the queue message got lost. The search index still has a file the database deleted. A worker crashed halfway through a job and nobody knows which half.
Then, this year, MinIOA popular self-hosted storage server that speaks the Amazon S3 protocol.'s community edition stopped receiving development, security patches and binaries, and most of the replacements hand back a BucketThe top-level container for files in S3-style storage, like a drive that holds folders and files. API (application programming interface)The set of requests one program accepts from another. A web API is how apps, scripts and AI agents ask a service to read or change its data. and nothing else.
Copal is what I built instead, and it's now open source under Apache License 2.0A permissive open-source license. You can use, change and redistribute the code, including inside closed-source products and hosted services, as long as you keep the license and copyright notices..
Most of a file service is bookkeeping, and bookkeeping belongs in one TransactionA group of database changes that succeed or fail together, so the data is never left half updated. database.
Copal keeps file metadata, the search index, the processing journal, access grants, API keys and the audit trail in a single SurrealDBAn open-source database that stores ordinary tables, linked records and nested documents in one engine. Kayak, Copal and Antumbra are all built on it. database. There's no external queue and no workflow engine. Every extra component I could add would be one more thing to operate, and none of them would give me a correctness guarantee the database doesn't already provide.
The name fits the job. Copal is tree resin, the young form of amber, and it holds on to whatever lands in it.
What Copal is
Copal is a Self-hostedSoftware you run on your own machines or your own cloud account, instead of renting it as someone else's service. file service written in RustA programming language known for being fast and for catching whole classes of bugs before a program ever runs., and it ships as one binary. You put files in over RESTThe most common style of web API. Each kind of thing gets a web address, and you read or change it with standard HTTP requests., GraphQLA query language for APIs in which the caller asks for exactly the fields it wants and gets them back in one response., S3Amazon's cloud file storage, and the protocol for storing and fetching files that many other storage systems now speak., tusAn open protocol for uploads that can pause and resume, so a dropped connection doesn't mean starting a large upload over. or MCP (Model Context Protocol)The standard way AI coding agents such as Claude Code connect to outside tools and data.. Copal checks each upload against its DigestA fixed-length fingerprint computed from a file's bytes. Any change to the file changes the digest, so it proves the bytes are intact., stores the bytes once no matter how many times they arrive, keeps every version, runs each upload through a processing pipeline, pulls out the text, and makes it searchable. Then it serves the bytes back under per-file access rules.
The bytes live in a Content addressingNaming data by a cryptographic hash of its bytes instead of by where it is stored. Anyone can serve it, and anyone can check that it matches the name. Blob storeStorage for raw file contents, called blobs, kept separately from the information about them. behind one interface: the local filesystem, S3, GCSGoogle Cloud Storage, Google's equivalent of Amazon S3. or AzureMicrosoft's cloud platform. Its file storage service plays the same role as Amazon S3., optionally Encrypted at restStored on disk in encrypted form, so the raw storage is unreadable to anyone without the key.. Everything else lives in SurrealDB.
Running it
The default build carries the database engine inside it, so a local instance needs no database server, no ContainerA packaged program bundled with everything it needs to run, so it behaves the same on any machine. and no config file:
cargo build -p copal-server -p copal-cli
COPAL_BIND=127.0.0.1:8099 COPAL_AUTH_MODE=keys COPAL_ADMIN_TOKEN=local-admin-token COPAL_DB_URL="surrealkv://./data/local/db" COPAL_BLOB_ROOT=./data/local/blobs COPAL_BLOB_ENCRYPTION_KEY=$(printf 'a%.0s' {1..64}) ./target/debug/copal-serversurrealkv:// is SurrealDB running inside the process and
persisting to a directory. Point COPAL_DB_URL at ws://
instead and the same binary talks to a SurrealDB server. On
first boot it reconciles the SchemaThe declared shape of a database or an API: which tables or types exist, what fields they have, and how they can be looked up. against the empty database
and starts listening.
Mint a key for a TenantOne customer, team or workspace whose data is kept separate from everyone else's inside a shared service., create a record, and send it some bytes:
export COPAL_URL=http://127.0.0.1:8099 COPAL_ADMIN_TOKEN=local-admin-token
TOKEN=$(curl -s -X POST $COPAL_URL/v1/admin/tenants/acme/keys -H "x-copal-admin-token: $COPAL_ADMIN_TOKEN" -H 'content-type: application/json' -d '{"name":"local","scopes":["read","write"]}' | jq -r .token)
ID=$(curl -s -X POST $COPAL_URL/v1/files -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"path":"runbooks/rotation.txt","content_type":"text/plain"}' | jq -r .id)
curl -s -X PUT $COPAL_URL/v1/files/$ID/content -H "authorization: Bearer $TOKEN" --data-binary @rotation.txtWhen I ran that for this article, the upload answered like this (trimmed a little):
{
"id": "01m3abz35r212xt2tfkva87202",
"path": "runbooks/rotation.txt",
"state": "scanning",
"size": 252,
"digest": "a0ac806c5553ad3ea9bd1fe18800d269f262501eb5af77dfe35f93ecc9ac5ce8",
"version_count": 1
}A moment later the same record read back as ready, with the
pipeline's findings written into its metadata:
"state": "ready",
"metadata": {
"processing": {
"sniffed_type": "text/plain",
"type_matches": true,
"extracted": true,
"passages": 1,
"verdict": "clean"
}
}The key that came back looks like ck1.<id>.<secret>, and
Copal only stores a hash of the secret. On disk, the bytes sit
at objects/a0/ac/<digest> and the file starts with CPE1,
the header of Copal's encrypted format, because I gave it a
key. The plaintext never touches the filesystem.
The text is searchable straight away:
curl -s "$COPAL_URL/v1/search?q=rotation" -H "authorization: Bearer $TOKEN"{
"mode": "lexical",
"items": [
{
"file": "01m3abz35r212xt2tfkva87202",
"passage": 0,
"excerpt": "Signing key rotation runbook.\nRotate the blob encryption key by...",
"matches": [[12, 20], [30, 36]]
}
],
"next_cursor": null
}The hits come back as passages with character offsets, so a
client can highlight them without re-searching. The mode
says lexical because I hadn't configured an EmbeddingA list of numbers that captures what a piece of text means, so text about similar things can be found by comparing the numbers. An embedder is the model that produces them.
service. Hybrid search is the default, and when Copal can't
run the semantic half it answers lexically and says so,
instead of quietly returning something different from what you
asked for.
The same file is already on every other face. The
contract-generated REST twin at /v1c, GraphQL, the MCP
EndpointOne specific web address a service answers requests on, such as the one that lists files. and the terminal client all see it:
curl -s -X POST $COPAL_URL/graphql -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"query":"{ files(limit: 3) { items { id path state } } }"}'
# {"data":{"files":{"items":[{"id":"01m3abz35r212xt2tfkva87202","path":"runbooks/rotation.txt","state":"ready"}]}}}
COPAL_TOKEN=$TOKEN ./target/debug/copalctl search "key rotation" --limit 3There's also an operator console at /admin/console. It's
plain server-rendered HTMLThe language web pages are written in. A page can include scripts that run in the browser of whoever opens it. with no external assets, so it
works on an Air-gappedPhysically cut off from other networks, including the internet. network. You log in with any username
and the admin token as the password.
What happens to a file
An upload is two requests. The first creates a record, which
starts as a draft. The second streams the bytes. Copal
hashes them on the way in, writes them to a staging key, and
then renames the staged object onto its content address. The
record flips to scanning in the same transaction that writes
the new version row and links the blob. Then the request
returns, and a worker picks up the rest.
The worker runs six steps in order. It sniffs the real type
from the first bytes and compares it with the declared one. It
checks the extension policy. If a ClamAVA free, open-source antivirus engine that can scan files as they arrive. daemon is configured
it streams the content through it. It extracts the text and
splits it into overlapping passages. If an embedding service
is configured it embeds them. Finally it flips the record to
ready or quarantined and writes the verdict into
metadata.processing, which is what you saw above.
Some of that is deliberately done elsewhere. Copal decodes
text and JSONA plain-text format for structured data, built from named fields and lists, that almost every programming language can read and write. itself and carries no document ParserA program that reads text or source code and works out its structure. Parsers tend to break when the format they expect changes.. PDF,
Office formats and OCR (optical character recognition)Turning pictures of text, such as scanned pages, into real text a computer can search. go to an extractor service with an API
shaped like Apache TikaAn open-source tool that pulls the text out of PDFs, Office documents and hundreds of other file formats.'s, because those parsers are large,
fast-moving, and historically a rich source of memory-safety
bugs. Embeddings come from any /v1/embeddings endpoint
shaped like OpenAIThe company behind ChatGPT. Its API format is widely copied, so many tools accept any service shaped like it.'s. Copal runs no models, because InferenceRunning a trained AI model to get an answer, as opposed to training it.
means WeightsThe learned numbers inside an AI model. Training adjusts them, and answering a question reads them., a runtime and hardware assumptions that have no
business inside a storage service.
One rule in that diagram matters more than it looks. Copal serves a file's bytes when the record has a digest and isn't quarantined. The lifecycle state alone never decides. A re-upload in flight doesn't blank out the version that's already there, and a re-upload that fails doesn't take the last good version offline. Tying serving to state would turn every transition into a small outage.
When I uploaded the same bytes a second time, the record went
straight to ready as version 2. The digest matched content
whose verdict the pipeline had already recorded, so there was
nothing to redo, and both versions point at the same blob on
disk.
The journal is the queue
The pipeline doesn't need a queue because it runs from a
journal in the same database. A run is a row, each step
attempt is a row, and a unique index on
(run_key, step_key, attempt) means that execution is
at-least-once but every step is recorded exactly once. If a
worker dies, its LeaseA time-limited claim on a piece of work. If the worker holding it dies, the claim runs out and another worker can take the work. expires, the run goes back to pending,
and the next worker replays the journal, skipping every step
that already finished. A crash costs one lease timeout.
The rest of the state handling follows the same habit. Every
state change carries its preconditions in the WHERE clause
of the UPDATE that performs it, and an empty result means
the caller lost the race. That makes the database the one
place where concurrent writes get ordered, however many Copal
instances are running. Every state that can get stuck has an
automatic way out:
| Stuck on | The way out |
|---|---|
| An upload claimed by an instance that died | The lease expires and a sweep marks the file failed, which is retryable |
| A run claimed by a worker that died | The lease expires, the run returns to pending, replay skips finished steps |
A file in scanning with no run behind it |
The stale-scan sweep fails it after a timeout |
| Staged bytes nobody finished | The staging sweep deletes them once they're old enough |
| Blobs nothing points at | Garbage collection recounts from live links, waits out a grace period, then collects |
Garbage collectionAutomatically finding and deleting data that nothing refers to anymore. never trusts a counter. It recomputes a blob's references from the live records every time, because counters that get incremented and decremented drift under crashes and races, and a count derived fresh is correct each time it's read.
Custom processing uses the same machinery. A workflow is a list of named activities, each an async function from JSON to JSON, and each step's output becomes the next step's input:
let registry = FlowRegistry::new()
.activity("resize", |input| async move {
// input is the previous step's output (or the run input)
Ok(serde_json::json!({ "resized": true, "source": input }))
})
.workflow("thumbnail", &["resize"], 3);
let state = AppState::new(store, blobs).with_flow(registry);Activities have to be IdempotentSafe to run more than once: doing it twice has the same effect as doing it once., since a replay after a crash
can run one again. For heavier work there's an external
transformer seam. The repo ships an ffmpegThe standard open-source tool for converting, cutting and inspecting video and audio files. example, a small
PythonA popular, easy-to-read programming language used for everything from small scripts to data science. server with thumbnail, audio, preview and probe
recipes. Its probe output walks the normal pipeline and gets
indexed like anything else, which means a search for h264
finds videos by CodecThe format audio or video is compressed in, such as H.264 for video..
Search
Search runs over passages rather than whole documents. Lexical searchSearch by the exact words in the text, as opposed to search by meaning. search is BM25A classic formula for ranking text by the words it shares with a query, giving rare words more weight than common ones., and Copal does the scoring itself. Semantic search uses an HNSWHierarchical navigable small world, a popular kind of vector index that finds close matches quickly by hopping through layered graphs of neighbors. index over the passage embeddings, stored at Half and full precisionHow many bits each number in a model uses. Half precision uses 16 bits instead of 32, halving the memory at a small cost in accuracy.. Hybrid, the default, fuses the two rankings with Reciprocal rank fusionA way to merge two ranked lists into one by rewarding items that rank near the top of either list., and an optional RerankerA second model that reads each search result next to the question and reorders the results by how well they actually answer it. can reorder the top results. Filters on where a file sits in its path and on its content type apply inside the database on both halves, so a filtered search never ranks passages it would only throw away.
GET /v1/search?q=signing+key&mode=hybrid&prefix=runbooks/&limit=20
GET /v1/search?q=inspection&facets=content_type,access
GET /v1/files/{id}/textEmbeddings look after themselves. When the configured model changes, a background pass re-embeds every passage that still carries the old model's vectors, and the Vector indexAn index over embeddings that finds the stored items closest in meaning to a query, instead of matching exact words. rebuilds itself when the dimension changes.
Sharing without handing out keys
Most file services hand out Signed URLA link that carries its own permission, so whoever has it can fetch one file without logging in, until it expires., which are awkward to
revoke because nothing on the server remembers issuing them.
Copal's grants are rows instead. A grant token is
cg1.<id>.<secret>, and the database keeps only a hash of the
secret, so there's no signing key to rotate or leak. Revoking
a grant is one write, and use counts are enforced atomically
as part of serving.
In my run I issued a grant with "max_uses": 1. The first
request got the file. The second got this:
{"error":{"kind":"not_found","message":"not found: unknown or unusable grant"}}Every refusal looks the same, whether the token is malformed, unknown, revoked, expired or used up, so a caller can't probe for which grants exist.
The same idea runs the other way for uploads. Your backend
creates the record, decides its path and access level, and
asks for an upload grant. The browser or phone then PUTs
straight to Copal with nothing but that URLA web address, such as https://example.com/report.pdf.. Upload grants are
single-use and capped at 24 hours, and the bytes go through
the same size limit, quota and pipeline as any other upload.
For a CDN (content delivery network)A company that keeps copies of websites on servers around the world so pages load quickly. Many large sites depend on the same few CDNs. there's a second family, cg2 edge tokens. They're
stateless HMACA signature made with a shared secret key, so anyone holding the same key can check that a message wasn't forged or changed. tokens that an edge worker holding the tenant's
edge secret can verify without calling back to the origin. The
trade is that you can't revoke one at a time, so they're meant
for short lifetimes.
Every file also carries an access level. public files serve
anonymously and tell caches to keep them for a year. private
and tenant files serve only to their tenant. grant files
refuse direct download for everyone, the owner included, so
the only way to their bytes is through a URL someone
deliberately issued.
One contract behind the API
Copal's API is declared once, as a
Kayak Contract (Kayak)In Kayak, the single checked-in declaration of everything an API offers, from which the documents, the clients and the live server are generated..
Kayak checks that declaration against the real indexes and
generates the OpenAPIA standard format for describing a REST API. Tools read an OpenAPI document to produce documentation, test tools and client code. document, the GraphQL SDL (Schema Definition Language)The plain-text format that describes a GraphQL API: its types, their fields, and the queries and changes it accepts., the MCP
manifest and clients in four languages. It also runs the
GraphQL, MCP, /v1c and console faces through one Dispatcher (Kayak)The part of Kayak's runtime that checks each request against the contract, applies limits and permissions, and only then calls your code.
that enforces scopes, rate budgets and Field guardA rule that decides, for each caller and each row, whether a field may be shown. Fields the caller may not see are removed from the answer.. The
hand-written /v1 routes stay canonical for the things a
generic router doesn't do yet, such as streaming bytes, ranges
and conditional requests. A parity test holds the two REST
faces to the same answers.
That matters most for agents. tools/list returned 22 tools
in my run, from files_list and file_create through
search and run_start. An agent's whole ingest loop is
three calls: file_create, file_issue_upload_url, and an
HTTPThe protocol web browsers, apps and servers use to request and send data over the web. PUT of the bytes to the grant URL. It goes through the
same scopes, budgets and guards as a script with the same key
would. What an agent may do is what its key may do. Keys can
also belong to principals of kind agent, which get their own
budgets and their own line in the audit trail.
Coming from MinIO
The S3 gatewayCopal's door for tools that speak the S3 protocol, so the same commands that work with Amazon S3 or MinIO work with Copal. takes the tools people already have. The bucket
is the tenant and the key is the path, so a migration is one
mirror:
mc alias set old http://minio:9000 <old key> <old secret>
mc alias set copal http://copal:9000 <access key> <secret>
mc mirror old/data copal/acme
mc diff old/data copal/acme # empty output means every object arrivedIn the recorded run in the migration guide, 80 MiB (mebibyte)1,048,576 bytes, a little more than a million. came across in 13 seconds, a second pass copied nothing, and an 80 MiB Multipart uploadUploading a large file as separately sent pieces that the server joins at the end, so one failed piece can be retried on its own. object round-tripped with a matching SHA-256A widely used cryptographic hash function that produces a 256-bit fingerprint of any data.. Every file that arrives that way gets scanned, DedupeStoring identical content only once, however many times it arrives or under however many names., versioned and indexed like any other upload, so the mirrored bucket can answer search queries, which a plain object store can't.
The Conformance suiteA set of tests that checks whether a system behaves the way a standard, or the existing tools that speak it, expect. runs stock clients against the S3 gateway, and anyone can run it on their own hardware:
| Client | Checks passed |
|---|---|
| aws CLI | 9 of 9 |
MinIO mc |
8 of 8 |
| rclone | 5 of 5 |
| MCP over curl | 4 of 4 |
The S3 gateway is a door onto Copal's storage and not a full
S3 implementation. Tagging, bucket policies, lifecycle rules,
ACL (access control list)A list attached to a file or resource that says exactly who may read or change it., the S3 versioning API and Object Lock all answer
501 NotImplemented. ETagA short tag a server attaches to a file so a client can tell whether the file has changed since it last looked. are SHA-256 rather than MD5An older hash function, still used as a quick file fingerprint but no longer considered secure., which
means aws s3 sync sees every object as changed.
What it drives
These are the jobs Copal is built for, and the ones I use it for or am building toward:
| Workload | What Copal gives it |
|---|---|
| A document store behind an AI memory system | Idempotent creates, versions per re-ingest, extracted text, provenance by digest |
| A searchable archive or RAG source | Passage-level lexical, semantic and hybrid search, with facets and an optional reranker |
| Uploads from browsers and phones | Single-use upload grants, so the client never holds a key |
| Media behind a CDN | Immutable public caching, edge tokens, renditions, an ffmpeg transformer seam |
| A MinIO replacement | An S3 gateway that stock tools accept, plus everything above for each migrated object |
| Compliance-sensitive storage | Retention, legal hold, per-residency encryption keys, an audit trail the database refuses to rewrite |
| Multi-tenant SaaS backends | Per-tenant keys, quotas and storage residencies, with tenancy optionally enforced a second time by the database |
| Event-driven integrations | Signed webhooks, a replayable change feed, GraphQL subscriptions |
| Tools for agents | 22 MCP tools under the same scopes and budgets as every other caller |
The first row is live today. Antumbra archives the original bytes of every document it ingests into Copal, and every passage it stores carries the Copal file id and digest it came from. The whole client is two calls:
let mut body = json!({
"path": document_path(&title, hash),
"content_type": "text/plain",
"idempotency_key": idempotency_key(hash),
"metadata": { "title": title, "workspace": workspace },
});
if let Some(src) = source {
body["metadata"]["source"] = json!(src);
}
let created = transport.post_json(&format!("{base}/v1/files"), &credential, &body)?;
let file_id = created
.get("id")
.and_then(Value::as_str)
.ok_or_else(|| AntumbraError::other("copal create response missing `id`"))?;
let uploaded = transport.put_bytes(
&format!("{base}/v1/files/{file_id}/content"),
&credential,
"text/plain",
content.as_bytes(),
)?;The idempotency key is derived from the workspace and title, so ingesting the same document again makes a new version of the same Copal file rather than a second file.
My own instance runs in k3sA lightweight version of Kubernetes, the system that runs and restarts containerized services across a group of machines. on my homelab as a single ReplicaOne running copy of a service. Several replicas can share the load or take over if one fails. with a 2 GiB (gibibyte)1,073,741,824 bytes, a little more than a billion. memory limit, against its own SurrealDB. Its embeddings come from a Jetson Orin NanoA small NVIDIA computer with a built-in GPU, made for running AI models on very little power. running llama.cppAn open-source program for running AI models efficiently on ordinary hardware, from laptops to small boards. on the same network.
It scales down further than that and up a little past it. The same binary runs in three shapes:
The two-instance shape is tested, not just drawn.
COPAL_HA=1 ./conformance/run.sh runs the whole conformance
suite through nginxA widely used web server, often placed in front of other services to spread requests between them. Round-robinHanding each new request to the next server in turn. with no Sticky sessionsSending every request from one user to the same server. A service that doesn't need them lets any server answer any request.. The
instances need a shared blob root, the same keys and one
database. Sweeps elect a leader through a database lease, and
rate budgets can be shared across the fleet through the
database too.
Some numbers
These come from my own machine: an i9-12900KSA high-end Intel desktop processor. with 64 GiB of RAMA computer's working memory, which holds whatever it is actively using. It is much faster than disk and is cleared when the power goes off. and NVMeA fast kind of solid-state storage that connects directly to the computer's processor. storage, with in-memory metadata, filesystem blobs, and encryption turned on. They describe the shape of the system rather than a tuned deployment's ceiling.
| Measurement | 512 MiB | 1 GiB | 2 GiB |
|---|---|---|---|
| Streamed single PUT | 1.6 s | 3.4 s | 6.4 s |
| Ranged 1 MiB read, start / middle / end | 2 / 2 / 1 ms | 2 / 3 / 2 ms | 2 / 2 / 2 ms |
| Full sequential read | 0.9 s | 1.8 s | 3.7 s |
| Peak memory, PUT and every read | 49 MiB | 49 MiB | 50 MiB |
Sealed uploads run at about 300 MiB/s, against 515 MiB/s for the same 1 GiB upload unencrypted. Ranged readReading just part of a file instead of the whole thing, which is how a video player jumps to the middle of a video. into an encrypted object stay around 2 ms (millisecond)A thousandth of a second. wherever they land, because the format seals in 64 KiB (kibibyte)1,024 bytes, a little more than a thousand. frames and a read only opens the frames it needs. Memory stays flat as files grow, with one exception: the malware scan pass holds about three times the object, which is worth knowing before you send ClamAV a 2 GiB video.
The embedded engine is also quick. A point read of one file record takes about 109 µs (microsecond)A millionth of a second. at the median when SurrealDB runs inside the process, against about 606 µs over a WebSocketA long-lived, two-way connection between a program and a server, so messages can flow either way without a new request each time. to a server on the same host.
What it won't do
It won't parse documents in-process or run models. Both are seams to services you run next to it.
It isn't a full S3 implementation. The S3 gateway covers what migration and everyday sync tools need, and stops there.
It won't render content that can run script. HTML, SVGA web image format for drawings, written as text. Because it can also carry scripts, services treat it as potentially active content. and XMLA text format for structured data built from nested tags, older and wordier than JSON. are always served as attachments.
It doesn't do multi-region writes. Replication is designed but not built, and for now a second region means a second deployment.
It keeps no write-ahead archive of its own, so point-in-time recovery comes from backing up the database and the blob store.
It doesn't have its own operator sign-in. Operator identity comes from a ProxyA server that sits in front of another and passes requests through to it, often handling tasks such as login on its behalf. in front of it, such as oauth2-proxy, Authelia or Cloudflare Access.
Current state of Copal
Copal hasn't cut a 0.1.0 release. CI (continuous integration)Automated checks that build and test every proposed change before it is merged, so problems are caught before they ship. publishes a container
image to ghcr.io/oneiriq/copal on every push to main, and
the code is what I run.
The defaults are open for development on purpose. The default
auth mode trusts an x-copal-tenant header, and it flips to
API keys at 1.0. Before exposing an instance, turn on keys
mode with an admin token, put the admin surface on its own
listener, terminate TLS (Transport Layer Security)The encryption that protects web traffic, shown by the padlock in your browser's address bar. in front of it, and give it a scoped
database user. The operations guide walks through all of it.
The generated client SDK (software development kit)A ready-made library that makes a service easy to use from one programming language. exist in Rust, TypeScriptJavaScript with type annotations added, widely used for web apps and servers. The types catch many mistakes before the code runs., Python
and Go, but they aren't published to any registry yet. They
authenticate with the development tenant header only, and they
don't move bytes. Uploads and downloads are plain HTTP, or a
PUT to an upload grant.
Search has two limits worth knowing. BM25 ranks a bounded window of matches, so a query that matches more than a few hundred passages ranks that window rather than every match. The search cursor is best-effort, because rankings shift as content changes.
What I lean on is the test suite. The whole service integration-tests against an in-memory SurrealDB engine with no containers involved, across 44 integration test files in the server alone, alongside the conformance suite and the BenchmarkA fixed set of test questions or tasks used to measure a system and compare it with others. above.
Getting it
Build from source with the commands in Running it, or pull the container image. The development loop is three commands:
docker compose up -d # SurrealDB v3, for the server-backed shape
cargo test # unit and in-memory round-trip tests, no server needed
cargo run -p copal-serverCopal is licensed under Apache-2.0.
Retrospect
The part of Copal I'm happiest with is how little there is to run. A file's bytes are in one place, and everything anyone knows about them (versions, verdicts, passages, grants, events, the audit trail) is in one database that can answer for all of it in a single transaction. When something goes wrong, there's one place to look. When something gets stuck, there's a sweep that knows the way out. It holds on to what you give it, and it can tell you what it's holding.
