Dear advisory board members,
Unfortunately, we were unable to record the presentation from the last advisory board meeting. To ensure everyone has access to the information, I am sharing a link to the presentation below, along with a summary of the key points and discussion highlights from the Q&A session.
Presentation incl. video recordings of sequence search in the UI: https://gbif.box.com/s/y1vvhfkr5wsq66foxywltd0j5yznzkst. You will also receive a separate email with a share link.
Discussion highlights:
Q: Is there a reliable way exclude metabarcoding data from searches? basisOfRecord = PreservedSpecimen?
A: We're going to be categorising the sources of data, and have a vocabulary underway: https://registry.gbif.org/vocabulary/DatasetCategory/concepts. You'll be able to filter in/out data originating from datasets on any of those categories.
Q: Are there any thoughts or concerns about the DSI language going on and ABS?
A: We have draft updates to the data user and data publisher agreement to include information on DSI. Data publishers will be required to share either country or countryCode when sharing sequences and when they know the location's origin. The requirements and recommendations will also be reflected in our technical documentation and DNA-derived and MDT guides. A news item has been drafted for a summary of the changes, which we plan to share at the next advisory board meeting. Next week, we are consulting with our node community and expect all information to go out before the GBIF Governing Board 33 in late September and the CBD-COP17 in Armenia in October.
Q: Will data providers be able to supply md5 hashes or only raw sequences that need to be sanitised?
A: We could promote a common approach to cleaning the sequences, so the MD5s will be comparable. However, the current recommendation<https://docs.gbif.org/publishing-dna-derived-data/en/#mapping-metabarcoding…> is that publishers share their own md5 in the TaxonID field, i.e. "ASV:7bdb57487bee022ba30c03c3e7ca50e1", and GBIF will populate the nucleotideSequenceID field with the md5 of the sanitised sequence, which would be identical to the TaxonID value if no cleaning was needed. Both fields will be searchable.
The next meeting will be in early October, and you will receive an invitation in August.
We look forward to connecting again in the autumn. In the meantime, we wish you all a very pleasant summer.
Kind regards
Cecilie
[cid:image001.png@01DD0EBA.8F305C00]<https://gbif.link/25>
Cecilie Svenningsen, PhD
Data Project Manager
GBIF: the Global Biodiversity Information Facility
Universitetsparken 15 | DK-2100 Copenhagen | DENMARK
GBIF.org<https://www.gbif.org/> | Mastodon<https://biodiversity.social/@gbif> | Bluesky<https://bsky.app/profile/gbif.org> | LinkedIn<https://www.linkedin.com/company/gbif/> | Facebook<https://www.facebook.com/gbifnews/>