# Cohort manifest content

**URL:** <https://discourse.canceridc.dev/t/cohort-manifest-content/37>\
**Category:** Developers\
**Tags:** portal\
**Created:** [August 10, 2020, 8:50pm UTC](https://discourse.canceridc.dev/t/cohort-manifest-content/37 "2020-08-10T20:50:12Z")\
**Posts on this page:** 7\
**Page:** 2

<div class="post-metadata">

**Author:** ![spaquett](https://sea1.discourse-cdn.com/flex015/user_avatar/discourse.canceridc.dev/spaquett/32/11_2.png) [@spaquett](https://discourse.canceridc.dev/u/spaquett)\
**Post date:** [August 14, 2020, 4:03pm UTC](https://discourse.canceridc.dev/t/cohort-manifest-content/37/21 "2020-08-14T16:03:24Z")

</div>

Since we don’t currently have infrastructure in place for tags, those would have to be post-MVP. A text description field is already available, and could be included.

---

<div class="post-metadata">

**Author:** ![fedorov](https://sea1.discourse-cdn.com/flex015/user_avatar/discourse.canceridc.dev/fedorov/32/3_2.png) [@fedorov](https://discourse.canceridc.dev/u/fedorov)\
**Post date:** [August 21, 2020, 1:30pm UTC](https://discourse.canceridc.dev/t/cohort-manifest-content/37/22 "2020-08-21T13:30:27Z")

</div>

@spaquett I think we should also include the DOI. I don’t know where it will be in the final layout of the tables, but at the moment this would be the `Source_DOI` from `idc-dev-etl:idc_tcia_views_mvp_wave0.dicom_all`.

cc: @bill.clifford

---

<div class="post-metadata">

**Author:** ![fedorov](https://sea1.discourse-cdn.com/flex015/user_avatar/discourse.canceridc.dev/fedorov/32/3_2.png) [@fedorov](https://discourse.canceridc.dev/u/fedorov)\
**Post date:** [August 21, 2020, 9:08pm UTC](https://discourse.canceridc.dev/t/cohort-manifest-content/37/23 "2020-08-21T21:08:29Z")

</div>

A post was split to a new topic: [TCIA manifest for the IDC cohorts](https://discourse.canceridc.dev/t/tcia-manifest-for-the-idc-cohorts/75)

---

<div class="post-metadata">

**Author:** ![fedorov](https://sea1.discourse-cdn.com/flex015/user_avatar/discourse.canceridc.dev/fedorov/32/3_2.png) [@fedorov](https://discourse.canceridc.dev/u/fedorov)\
**Post date:** [August 24, 2020, 5:23pm UTC](https://discourse.canceridc.dev/t/cohort-manifest-content/37/24 "2020-08-24T17:23:08Z")

</div>

@spaquett I now think that we should include full `gs://` URL to the individual files. This would be most convenient for the users that would want to replicate the data on the VM, and filelist can be piped to `gsutil`.

---

<div class="post-metadata">

**Author:** ![bill.clifford](https://avatars.discourse-cdn.com/v4/letter/b/8dc957/32.png) [@bill.clifford](https://discourse.canceridc.dev/u/bill.clifford)\
**Post date:** [September 9, 2020, 7:31pm UTC](https://discourse.canceridc.dev/t/cohort-manifest-content/37/25 "2020-09-09T19:31:16Z")

</div>

Finally getting around to this thread…  
To me the basic question is what are manifest use cases? In particular, is a manifest something that a user will use once and then throw away? Or is a manifest something that a user will squirrel away?  
Remember that the user can always get a new copy of the manifest from IDC.

A throw away manifest would be based on current (versioned) URLs, while a long lived manifest should be based on DOIs. I think it is debatable whether a manifest should include both. Perhaps best would be to only provide DOI based manifests, but provide a tool for easily obtaining the corresponding files. In this way there is no chance of a manifest becoming “stale”.

A DOI based manifest could also be more concise than a GCS URL based manifest. Note that a GCS URL based manifest must include a URL for every instance in the cohort because we can only version GCS entities at the instance level (without data duplication…no symlinks in GCS). However, a DRS DOI representing a series or study will be version specific so a cohort based on DOIs can capture the data in a cohort more efficiently.

---

<div class="post-metadata">

**Author:** ![fedorov](https://sea1.discourse-cdn.com/flex015/user_avatar/discourse.canceridc.dev/fedorov/32/3_2.png) [@fedorov](https://discourse.canceridc.dev/u/fedorov)\
**Post date:** [September 9, 2020, 8:09pm UTC](https://discourse.canceridc.dev/t/cohort-manifest-content/37/26 "2020-09-09T20:09:04Z")

</div>

I think for the MVP we should address the most immediate need: efficiently get the data corresponding to the manifest onto a VM. I think for the MVP we should use GCS URLs in the manifest.

Are there tools that would be equivalent in performance to parallel download with `gsutil -m` but work with DRS DOIs?

What is used in the manifests in ISB-CGC? What are the manifest use cases that proved to be important to the ISB-CGC users?

---

<div class="post-metadata">

**Author:** ![bill.clifford](https://avatars.discourse-cdn.com/v4/letter/b/8dc957/32.png) [@bill.clifford](https://discourse.canceridc.dev/u/bill.clifford)\
**Post date:** [September 9, 2020, 8:17pm UTC](https://discourse.canceridc.dev/t/cohort-manifest-content/37/27 "2020-09-09T20:17:59Z")

</div>

Yes, for the MVP, we should go with GCS URLs. I recently added a column of GCS URLs (with version suffix) to dicom\_all.

I am not aware of any tools that work directly with DRS objects, probably something we have to create.

[Previous page](https://discourse.canceridc.dev/t/cohort-manifest-content/37.md?page=1)
