Dataset
Usage
Dataset is the primary unit of publication and citation. It contains its own metadata (public or private) and a list of Arbogen IDs, which act as pointers to sequence records rather than embedding the sequences themselves. Datasets can be shared across multiple Working Groups and, once stable, can obtain a DOI for citation.
Visibility: Dataset visibility is public by design: the dataset and the Arbogen IDs it references are discoverable. However, Arbogen IDs only point to sequences: they do not include FASTA data or sensitive metadata.
- If a sequence is public, it can be accessed or downloaded via its Arbogen ID.
- If a sequence is private, its ID may appear in a dataset, but the sequence and metadata remain inaccessible unless explicitly shared by the owner (see Data Sharing).
Technical behavior
- Versioning: each validated modification (addition/removal) creates a new version of the dataset, linked to its predecessor.
- Provenance & traceability: links between versions ensure traceability.
- Access control: public visibility includes only Arbogen IDs, creation date, and description.
Categories:
- Mine: Datasets created by you.
- Shared: Datasets shared with you via Working Groups.
- Reference: Platform reference datasets.
- All: Registry of all datasets on the platform.
Actions:
- Share: Share a dataset to a Working Group.
- Download:
- Arbogen IDs (CSV)
- Metadata (CSV)
- FASTA sequences
- Full export packages
- Open permalink: Open a dataset detail page via its permalink.
- Delete: Only the Owner can delete a dataset, and only if it does not have a DOI. Deletion removes it from all Working Groups.
Creation:
- Click Create Dataset.
- Fill in:
- Name
- Public Description
- Private Description
- Once created, the dataset is empty. Use Manage Data to add sequences via Arbogen IDs and then, when appropriate, create a DOI.
Dataset Sharing and Ownership Constraints
When working within a Working Group, datasets and sequences can be shared with all group members. However, access to the underlying data depends on ownership.
Partial visibility when sharing datasets
A dataset may contain sequences owned by different users.
When a user shares a dataset:
- The dataset structure (including all contained Arbogen_IDs) is visible to all group members
- However, only the sequences owned by the sharing user are fully accessible (metadata and sequence)
- Sequences owned by other users remain private and inaccessible (only the Arbogen_ID is visible)
This means that group members may see the full dataset, but only have access to all Arbogen_IDs and a subset of its data (metadata and sequence).
Full access through re-sharing by the owner
If the dataset owner is part of the Working Group, they can re-share the dataset.
When the owner grants full access:
- The dataset is shared again under the owner’s authority.
- All sequences within the dataset become accessible (metadata and sequence).
- Previously restricted sequences are now visible to all group members.
This action effectively updates the dataset access scope for the entire group.
Key principle
Access to sequence data is always controlled by the sequence owner, even when sequences are part of a shared dataset.
Dataset sharing does not override individual ownership permissions unless explicitly re-shared by the owner.