Analysis metadata
Analysis metadata describes processed data files and connects them to the relevant study and samples. In EGA, an Analysis is a metadata object that links one or more files to the samples they describe and records how those files should be understood.
When to use an Analysis record
In the FEGA Sweden Submitter Portal, Analysis records are normally used for:
- reference-alignment files, such as BAM or CRAM files
- sequence-variation files, such as VCF files
- sensitive phenotype files that are linked to samples
Reference-alignment BAM or CRAM files may be paired with BAI or CRAI index files. An index file is optional and can be generated from the alignment file when needed.
Data types without a dedicated metadata category
An Analysis record can also be used for data types that are not explicitly supported by the EGA metadata model. This is a fallback rather than the preferred approach. For these files, select Sample Phenotype as the Analysis type in the portal, even when the files do not contain phenotype data. Contact the FEGA Sweden Helpdesk before preparing metadata so that we can help choose an appropriate file format, describe the files, and link them to the relevant study and samples.
For each Analysis record, you link the relevant study, sample or samples, and files. You can also link an Experiment record when one is available. This makes it clear which data belong to which samples and how the files relate to the study.
What to prepare
Before entering the metadata in the portal, prepare:
- a clear title and short description for each group of related analysis files
- the Analysis type: Reference Alignment, Sequence Variation, or Sample Phenotype
- the study and sample identifiers that the files should be linked to
- for alignments and variant files, the reference genome and a concise description of the method used to produce the files
- the files and, where applicable, their associated index files
For an alignment or variation analysis, the portal also asks for information such as the sequencing platform and experiment type. Keep a short record of the analysis pipeline, software versions, and important settings so that you can describe the work accurately when needed.
Phenotype data in an Analysis record
Phenotype information that cannot be shared openly can be submitted as analysis data. A Phenopacket is one possible structured format for phenotype data, but it is not required. The FEGA Sweden Helpdesk can advise on a suitable format for your data. For simple tabular phenotype data, a TSV file often works well.
Do not include personal identifiers or sensitive phenotype values in the title or description: these are public metadata. Instead, use the public description to explain at a high level what the protected file contains and how it relates to the samples. The phenotype file itself remains under controlled access.
Get help with the structure
The FEGA Sweden Helpdesk can help you decide whether files should be registered as an Analysis record and what information is needed. Contact us when your metadata is ready for review, before you click the Finalize button in the Submitter Portal.
For the current portal fields and examples, see the EGA Submitter Portal documentation.