Portal documentation · Swagger UI · Features · Architecture · Getting started · Pitfalls · Contributing
Note
Aruna v3 is now in public testing. You can try it out and share feedback through GitHub issues. Aruna v2 remains available on the v2 branch.
Aruna helps research organizations store, describe, share and reuse data while keeping control of their own infrastructure. Research data is often spread across universities, labs, archives and computing centers, each with its own storage systems and access rules. With Aruna, each organization runs a node that connects directly to other nodes. Researchers can find and work with data across institutions, with descriptions that help them understand what each dataset contains, how it was created and how it can be reused.
- Web portal: browse files, edit datasets, manage access and follow compute runs in the browser.
- Works with your existing tools: every node offers an S3-compatible interface, so common tools, scripts and workflow systems work without changes.
- Rich dataset descriptions: datasets are described with RO-Crate, a widely used standard that covers files, people, instruments, software and workflows. Crates reference other crates to connect datasets with their sources and the analyses that produced them.
- Quality checks: profiles specify which details a dataset description should include, and Aruna highlights anything missing.
- Search across nodes: find datasets on all connected nodes, limited to what you are allowed to see.
- Groups and permissions: control who can access your data, with permissions for individuals, groups and roles.
- Automatic merging: edits made on different nodes are combined automatically when the nodes reconnect, including edits made offline.
- Native Git for datasets: every dataset is also a Git repository in the ARC layout. Clone it, work on branches and push changes back; the dataset description stays in sync.
- Dataset history: see every version of a dataset, compare versions and merge draft changes in the portal.
- Publish with a DOI: send datasets to Zenodo or other InvenioRDM repositories and keep both in sync. Existing records can be imported, too.
- Compute jobs: run analysis containers on your own infrastructure from the portal or through the standard GA4GH TES interface.
- Interactive notebooks: write and run Jupyter notebooks in the portal, with direct access to your data.
- AI assistants: connect AI assistants through MCP so they can search, read and work with your data, with your permissions.
- Flexible storage: keep data on local disks or connect other storage systems. Buckets can combine local files, copies and references to files on other nodes.
- Encrypted buckets: a bucket can store new files encrypted, with keys that stay with the people you choose.
- Safe transfers: Aruna checks files when storing and copying them to detect data corruption early.
- Open standards: single sign-on with OIDC, data references with GA4GH DRS and metadata harvesting with OAI-PMH.
- Simple to deploy: run a single node as one program, or a cluster of nodes.
Institutions need to control where their data is stored and who can access it. Moving files manually takes time, and the context needed to understand them can get lost along the way. Aruna helps institutions share data while retaining control over its storage and access.
- Nodes: each organization runs its own node. The node decides where the organization's data lives and who may access it.
- Realms: nodes that trust each other form a realm, for example an institute, a consortium or a project. Being in the same realm does not give anyone access to data; access is always granted explicitly through groups, roles and permissions.
- Direct connections: nodes can connect to each other automatically, including from behind firewalls. No central server is needed, and a node keeps working when others are offline. Changes are shared again once the nodes can reach each other.
- Data and description together: files and their descriptions move together, so a dataset stays understandable wherever it is used. Crates reference other crates, forming a provenance graph that connects source data, analyses and results. These links help researchers trace how a result was produced and understand what they need to reproduce it.
The goal is to make research data FAIR: findable, accessible, interoperable and reusable, while each institution stays responsible for its own data.
Start with the public portal, or run Aruna yourself using one of the options below.
You can explore Aruna without installing anything:
- Open the public v3 portal and look around.
- Read the portal documentation for step-by-step guides.
- Developers can explore the REST API in the Swagger UI.
This starts three connected nodes on your machine, so you can see how nodes work together. You
need docker with Docker Compose and, for convenience, just.
just preview # three nodes, a login server and the web portal
just local-cluster # three nodes without the portalWhen everything is ready, the command prints the addresses of each node, test logins and an
admin token. Press Ctrl-C to stop the cluster, or run just stop if the terminal was closed.
To run one node yourself, for example to test it with your own login server:
cp .env.example .env
cargo run -p arunaBuilding from source needs the Rust version named in rust-toolchain.toml.
The node then serves the API documentation on http://127.0.0.1:3000/swagger-ui and the S3
interface on http://127.0.0.1:1337.
A group admin can turn on encryption for a bucket, in the portal or with
PUT /data/buckets/{bucket}/storage/encryption. The node then stores new files of that bucket
encrypted.
- Two modes. With
node_managed, the node can open the bucket key after a restart. Withvault_locked, a key holder must unlock the bucket in the portal; the key stays only in the node's memory, until it is locked again, a time limit ends, or the node restarts. - Key holders. The bucket creator, the group admins and users granted explicitly hold a sealed copy of the bucket key. Keep at least two holders, or one with a recovery key, so the data is not lost with one person's key.
- While locked. Uploads, listings and object info still work. Reading content, copying it to
another bucket and jobs that need it wait or are refused. S3 answers
403 AccessDeniedwith the headerx-aruna-bucket-locked: true. - Token credentials. A key holder can create an S3 credential that reads chosen buckets while
they are locked: list them in
encrypted_bucketsofPOST /access/credentials. Each bucket must be unlocked at that moment. The answer contains asession_tokenonce; the node does not keep it. Set it asaws_session_tokennext to the access key and secret, for example in~/.aws/credentials. The token only works in the signed request header, never in presigned URLs. It stops working when the credential is revoked, when its creator stops being a key holder, or after a key rotation.GET /data/buckets/{bucket}/storage/encryption/tokenslists the token credentials of a bucket. - S3 clients. Encrypted buckets report
AES256server-side encryption.PutBucketEncryptionwithAES256turns onnode_managedmode; turning encryption off is only possible through the REST API. Other encryption headers, such as KMS or customer keys, are refused. - What is not covered. Turning on encryption does not remove plaintext that existed before: backups, old disk blocks and copies on other systems stay as they were. Deleting a file is not secure erasure. Job and notebook working directories are removed when the job or session ends, but logs, reports and files written by your own code are kept as they are; only files written to an encrypted bucket are encrypted.
- Content fingerprints. For deduplication, the node keeps content hashes of encrypted objects. These hashes are extra plaintext fingerprints: they can show that two files are equal.
- Copies and sync to other nodes. A target bucket that encrypts stores every copy sealed to
its own key. When both buckets encrypt, the source node grants each encrypted file to the key of
the target bucket, so the target never needs the source key. While the source bucket is locked,
its copies wait and run after the next unlock; the sync status shows how many wait. The target
checks the content of such a copy when its own bucket is next unlocked. A target bucket without
encryption gets copies of an encrypted bucket only when the request sets
plaintext: trueand the requester holds the key of the source bucket.
By default, a realm lists itself in the global registry at https://registry.aruna-engine.org,
so the portals of other realms can find it.
- What it sends. A registration signed with the realm key, renewed every 6 hours. It holds the realm's public facts (realm id, name, description, public API and portal URL) and three counts: live datasets, groups and configured nodes. It holds no users, data or metadata.
- How it starts. When the realm has no federation settings, one management node creates them
on start: the name from
REALM_DESCRIPTION, the URLs fromAPI_PUBLIC_URLandPORTAL_PUBLIC_URL, the default registry and registration turned on. Existing settings are never changed. - Only public realms. No settings are created when one of the two URLs is missing or points
to a loopback or private address. Local tools (
just preview,just local-cluster, Compose) never register. - Opting out. Before the first start, set
FEDERATION_REGISTRY_URLto another registry, or to an empty value to turn the default off. Once the settings exist, a realm admin changes them withPUT /api/v1/system/realm/federation. Turning registration off sends up to three withdrawals; removing the registry sends nothing more, and the entry expires after 72 hours.
Most problems with a new node come from a few settings. These tips help you avoid them.
- Use your own keys. The example configuration contains demo keys that everybody can see. A
node refuses to start with them. Create your own realm and node keys (
REALM_*_KEYandNODE_*_KEY) before running a real node. - Set up a node right the first time. A node creates or joins its realm only on its very first
start: without
ONBOARDING_SECRETit creates a new realm, with it it joins an existing one. Changing the configuration afterwards does not move it to another realm. To start over, use a new, empty data directory. - Back up the whole data directory. The data directory (
STORAGE_PATH) also holds the node's identity, which protects stored secrets such as access tokens. A node restored without it cannot read these secrets anymore. - Migrate before upgrading. Stop the node, run
aruna-doctor migrateon its data directory, then start the new version. Running the migration twice does no harm. - Keep enough free disk space. Importing a large dataset briefly needs about twice its size.
- Decide how changes are saved. By default, Aruna favors speed: the last few changes can be
lost if the machine loses power. Set
ARUNA_FJALL_PERSIST_MODE=sync_allif that is not acceptable for you. - Keep notebook networks separate. Notebook sessions run in their own network so they can only
reach your data. Make sure this network (
ARUNA_COMPUTE_DOCKER_SESSION_SUBNET) does not overlap with other networks on the host. - Use Git the usual way. Log in to the Git repository of a dataset with an Aruna access token as the password, pull before you push, and store large files with Git LFS.
- Portal documentation: how to use Aruna in the browser.
- Swagger UI: the full REST API.
- Native Git guide: working with datasets as Git repositories.
- Repository structure: how the source code is organized.
Aruna is licensed under either of
- Apache License, Version 2.0 (LICENSE-APACHE or http://www.apache.org/licenses/LICENSE-2.0)
- MIT license (LICENSE-MIT or http://opensource.org/licenses/MIT)
at your option. Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in Aruna by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.
Found a bug or have an idea? Open an issue or send a pull request. Reports from the v3 public test phase help us understand what works and what needs attention. See the Contributor Guidelines and Code of Conduct before contributing.