Most of our work has resulted in scholarly publications. On this page you can review our publications to get an idea about our work.
Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run them… Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run themselves in Terra. The most common reason for this is that the pipeline
accesses proprietary data that users are not allowed to access directly, but that may be used as e.g. a reference panel
for imputation.
Supported pipelines
Current supported pipelines are:
Array Imputation with the All of Us + AnVIL Reference Panel
Architecture
Architecture Doc
Architecture Diagram
Development
This codebase is in initial development.
Requirements
Technical
This service is written in Java 17, and uses Postgres 15.
To run locally, you'll also need:
jq - install with brew install jq
Java 17 - can be installed manually or through IntelliJ which will do it for you when importing the project
Postgres 15 - multiple solutions here as long as you have a postgres instance running on localhost:5432 the local app will connect appropriately. Be sure to use Postgres 15 (as of Feb 2025, Postgres 17 did not work)
Download Postgres.app (recommended) from https://postgresapp.com/
Brew https://formulae.brew.sh/formula/postgresql@15
External Services
Terra services
Sam
Used to authn users connecting to the service and authz users for admin endpoints
Rawls
Used to handle workspace interactions
creating methods
data tables
workflow submission
Cromwell
Used through Rawls to run submissions
Thurloe
Used to send notification emails to users
Tech stack
Java 17 temurin
Postgres 15
Gradle - build automation tool
SonarQube - static code security and coverage
Trivy - security scanner for docker images
Jib - docker image builder for Java
Local development
To run locally:
Make sure you have the requirements installed from above. We recommend IntelliJ as an IDE.
Clone the repo (if you see broken inputs build the project to get the generated sources)
Spin up a local postgres instance (NOTE: use version 15)
Run the commands in scripts/postgres-init.sql in your local postgres instance. You will need to be authenticated to access GSM.
Run scripts/write-config.sh
Run ./gradlew bootRun to spin up the server.
Navigate to http://localhost:8080/#
If this is your first time deploying to any environment, be sure to use the admin endpoint /api/admin/v1/pipelines/{pipelineName}/{pipelineVersion} to set your pipeline's workspace id.
To run this endpoint, you need to be authenticated using your firecloud test account. A list of accounts that developers typically need is here. Further, a list of resources that are generally useful is stored here
This endpoint requires two parameters directly, and three in the message body:
pipelineName can be retrieved by querying the /api/pipelines/v1 endpoint.
pipelineVersion can also be retrieved from the /api/pipelines/v1 endpoint.
workspaceBillingProject is listed in the Teaspoons Resources document linked above
workspaceName is also listed in the Teaspoons Resources document, and can be found through the Terra UI workspace dashboard
wdlMethodVersion is found for the specific workflow as listed in the Terra UI page for workflows.
Back up local Postgres databases before testing/refactors
Before running local migrations/refactors, take backups of both local databases so you can restore quickly.
Defaults in this repo (see service/src/main/resources/application.yml and scripts/postgres-init.sql):
host 127.0.0.1, port 5432
pipelines_db user/pass: dbuser / dbpwd
teaspoons_stairway_db user/pass: stairwayuser / stairwaypwd
Backup and verify:
ts="$(date +%Y%m%d_%H%M%S)"
backup_dir="$HOME/teaspoons-db-backups/$ts"
mkdir -p "$backup_dir"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_dump -Fc -f "$backup_dir/pipelines_db.dump" pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_dump -Fc -f "$backup_dir/teaspoons_stairway_db.dump" teaspoons_stairway_db
pg_restore -l "$backup_dir/pipelines_db.dump" | head
pg_restore -l "$backup_dir/teaspoons_stairway_db.dump" | head
echo "Backups written to: $backup_dir"
Restore later (replace <admin_password>):
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O dbuser pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_restore --clean --if-exists --no-owner -d pipelines_db "$backup_dir/pipelines_db.dump"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O stairwayuser teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_restore --clean --if-exists --no-owner -d teaspoons_stairway_db "$backup_dir/teaspoons_stairway_db.dump"
Local development with the UI
When running terra-ui locally against a local teaspoons backend, CORS-related errors can arise. To get around this, run the following command to copy a configuration file that allows requests from localhost:
./scripts/local-dev/copy_web_config.sh
Note that this file at the destination path (next to App.java) is ignored via .gitignore, since it should not be used in deployed environments.
Local development with debugging
If using Intellij (only IDE we use on the team), you can run the server with a debugger. Follow
the steps above but instead of running ./gradlew bootRun to spin up the server, you can run
(debug) the App.java class through intellij and set breakpoints in the code. Be sure to set the
GOOGLE_APPLICATION_CREDENTIALS=config/teaspoons-sa.json in the Run/Debug configuration Environment Variables.
Testing the CLI locally
If you make changes to openapi.yml, you should test the CLI locally.
To create the autogenerated Python client files locally, run
./gradlew :python-client:openApiGenerate
The files will be generated in python-client/generated and are ignored from being checked into the repo.
(Note: the unqualified ./gradlew openApiGenerate now regenerates all four codegen modules —
python-client, rawls-client, client, and service — so qualify the task when you only want the Python client.)
To test with the CLI, follow the instructions in the CLI repo: DataBiosphere/terra-scientific-pipelines-service-cli.
Running Tests Locally
Run ./gradlew service:test to run tests
Note: If you encounter errors indicating a failure to load the ApplicationContext due to an error while preparing a database cluster caused by a missing Docker environment,
this may be related to newer Docker versions (for example, 29.0.0 and above). To resolve this issue, override the
Docker API version in the $HOME/.docker-java.properties file. If the file does not already exist, create it and add the following line:
api.version=1.44
If the file mentioned already exists with above line, and the tests are still failing in the same way, try restarting Docker.
Running Linter Locally
Run ./gradlew spotlessCheck to run linter checks
Run ./gradlew :service:spotlessApply to apply fix any issues the linter finds
(Optional) Install pre-commit hooks
[scripts/git-hooks/pre-commit] has been provided to help ensure all submitted changes are formatted correctly. To install all hooks in [scripts/git-hooks], run:
git config core.hooksPath scripts/git-hooks
Running SonarQube locally
SonarQube is a static analysis code that scans code for a wide
range of issues, including maintainability and possible bugs. Get more information from
DSP SonarQube Docs
If you get a build failure due to
SonarQube and want to debug the problem locally, you need to get the sonar token from GSM
before running the gradle task.
export SONAR_TOKEN=$(gcloud secrets versions access latest --project="broad-dsde-dev" --secret="teaspoons-sonarcloud" | jq '.sonar_token')
./gradlew sonarqube
Running this task produces no output unless your project has errors. To
generate a report, run using --info:
./gradlew sonarqube --info
Connecting to the database
To connect to the Teaspoons database, we have a script in dsp-scripts that
does all the setup for you. Clone that repo and make sure you're either on Broad Internal wifi or connected
to the VPN. Then run the following command:
./db/psql-connect.sh dev teaspoons
Deploying to dev
Upon merging to main, the dev environment will be automatically deployed via the GitHub Action Bump, Tag, Publish, and Deploy
(that workflow is defined here).
The two tasks report-to-sherlock and set-version-in-dev will prompt Sherlock to deploy the new version to dev.
You can check the status of the deployment in Beehive and in
ArgoCD.
For more information about deployment to dev, check out DevOps' excellent documentation.
Tracing
We use OpenTelemetry for tracing, so that every request has a tracing span that can
be viewed in Google Cloud Trace.
See this DSP blog post for more info.
Running the BEE end-to-end tests
The end-to-end test that runs against a BEE is specified in .github/workflows/run-bee-e2e-tests.yaml. It calls the workflow defined
in the terra-github-workflows repo.
The end-to-end test is automatically run nightly on the dev environment.
To run the test against a specific feature branch:
Grab the image tag for your feature branch.
If you've opened a PR, you can find the image tag as follows:
go to the Bump, Tag, Publish, and Deploy workflow that's triggered each time you push to your branch
From there, go to the tag-publish-docker-deploy task
Expand the "Construct docker image name and tag" step
The first line should contain the image tag, something like "0.0.81-6761487".
Navigate to the e2e-test GHA workflow
Click on the "Run workflow" button and select your branch from the dropdown
Enter the image tag from step 1 in the "Custom image tag" field
If you've updated the end-to-end test in the dsp-resuable-workflows repo, enter either a commit hash or your git
branch name. If you don't need to change the test, leave the default as main.
Click the green "Run workflow" button.
Python clients
We publish a "thin", auto-generated Python client that wraps the Teaspoons APIs. This client is published to
PyPi and can be installed with
pip install teaspoons_client, although this is not meant to be user-facing. The thin api client is generated from
the OpenAPI spec in the openapi directory.
Publishing occurs automatically when a new version of the service is deployed, via the
release-python-client GHA.
We also have a user-facing, "thick" CLI whose code lives in a separate repository: DataBiosphere/terra-scientific-pipelines-service-cli. This repository contains the data and workflows for the publication "GalaxyMemMD: an end-to-end Galaxy toolset for reproducible membrane embedding, molecular dynamics simulation, and analysis of … This repository contains the data and workflows for the publication "GalaxyMemMD: an end-to-end Galaxy toolset for reproducible membrane embedding, molecular dynamics simulation, and analysis of membrane proteins". GalaxyMemMD provides a set of Galaxy tools and workflows that take a membrane protein from its initial structure through membrane embedding, molecular dynamics (MD) simulation, and trajectory analysis. This record archives the inputs, outputs and Galaxy workflows needed to reproduce the results in the paper. Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run them… Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run themselves in Terra. The most common reason for this is that the pipeline
accesses proprietary data that users are not allowed to access directly, but that may be used as e.g. a reference panel
for imputation.
Supported pipelines
Current supported pipelines are:
Array Imputation with the All of Us + AnVIL Reference Panel
Architecture
Architecture Doc
Architecture Diagram
Development
This codebase is in initial development.
Requirements
Technical
This service is written in Java 17, and uses Postgres 15.
To run locally, you'll also need:
jq - install with brew install jq
Java 17 - can be installed manually or through IntelliJ which will do it for you when importing the project
Postgres 15 - multiple solutions here as long as you have a postgres instance running on localhost:5432 the local app will connect appropriately. Be sure to use Postgres 15 (as of Feb 2025, Postgres 17 did not work)
Download Postgres.app (recommended) from https://postgresapp.com/
Brew https://formulae.brew.sh/formula/postgresql@15
External Services
Terra services
Sam
Used to authn users connecting to the service and authz users for admin endpoints
Rawls
Used to handle workspace interactions
creating methods
data tables
workflow submission
Cromwell
Used through Rawls to run submissions
Thurloe
Used to send notification emails to users
Tech stack
Java 17 temurin
Postgres 15
Gradle - build automation tool
SonarQube - static code security and coverage
Trivy - security scanner for docker images
Jib - docker image builder for Java
Local development
To run locally:
Make sure you have the requirements installed from above. We recommend IntelliJ as an IDE.
Clone the repo (if you see broken inputs build the project to get the generated sources)
Spin up a local postgres instance (NOTE: use version 15)
Run the commands in scripts/postgres-init.sql in your local postgres instance. You will need to be authenticated to access GSM.
Run scripts/write-config.sh
Run ./gradlew bootRun to spin up the server.
Navigate to http://localhost:8080/#
If this is your first time deploying to any environment, be sure to use the admin endpoint /api/admin/v1/pipelines/{pipelineName}/{pipelineVersion} to set your pipeline's workspace id.
To run this endpoint, you need to be authenticated using your firecloud test account. A list of accounts that developers typically need is here. Further, a list of resources that are generally useful is stored here
This endpoint requires two parameters directly, and three in the message body:
pipelineName can be retrieved by querying the /api/pipelines/v1 endpoint.
pipelineVersion can also be retrieved from the /api/pipelines/v1 endpoint.
workspaceBillingProject is listed in the Teaspoons Resources document linked above
workspaceName is also listed in the Teaspoons Resources document, and can be found through the Terra UI workspace dashboard
wdlMethodVersion is found for the specific workflow as listed in the Terra UI page for workflows.
Back up local Postgres databases before testing/refactors
Before running local migrations/refactors, take backups of both local databases so you can restore quickly.
Defaults in this repo (see service/src/main/resources/application.yml and scripts/postgres-init.sql):
host 127.0.0.1, port 5432
pipelines_db user/pass: dbuser / dbpwd
teaspoons_stairway_db user/pass: stairwayuser / stairwaypwd
Backup and verify:
ts="$(date +%Y%m%d_%H%M%S)"
backup_dir="$HOME/teaspoons-db-backups/$ts"
mkdir -p "$backup_dir"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_dump -Fc -f "$backup_dir/pipelines_db.dump" pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_dump -Fc -f "$backup_dir/teaspoons_stairway_db.dump" teaspoons_stairway_db
pg_restore -l "$backup_dir/pipelines_db.dump" | head
pg_restore -l "$backup_dir/teaspoons_stairway_db.dump" | head
echo "Backups written to: $backup_dir"
Restore later (replace <admin_password>):
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O dbuser pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_restore --clean --if-exists --no-owner -d pipelines_db "$backup_dir/pipelines_db.dump"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O stairwayuser teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_restore --clean --if-exists --no-owner -d teaspoons_stairway_db "$backup_dir/teaspoons_stairway_db.dump"
Local development with the UI
When running terra-ui locally against a local teaspoons backend, CORS-related errors can arise. To get around this, run the following command to copy a configuration file that allows requests from localhost:
./scripts/local-dev/copy_web_config.sh
Note that this file at the destination path (next to App.java) is ignored via .gitignore, since it should not be used in deployed environments.
Local development with debugging
If using Intellij (only IDE we use on the team), you can run the server with a debugger. Follow
the steps above but instead of running ./gradlew bootRun to spin up the server, you can run
(debug) the App.java class through intellij and set breakpoints in the code. Be sure to set the
GOOGLE_APPLICATION_CREDENTIALS=config/teaspoons-sa.json in the Run/Debug configuration Environment Variables.
Testing the CLI locally
If you make changes to openapi.yml, you should test the CLI locally.
To create the autogenerated Python client files locally, run
./gradlew :python-client:openApiGenerate
The files will be generated in python-client/generated and are ignored from being checked into the repo.
(Note: the unqualified ./gradlew openApiGenerate now regenerates all four codegen modules —
python-client, rawls-client, client, and service — so qualify the task when you only want the Python client.)
To test with the CLI, follow the instructions in the CLI repo: DataBiosphere/terra-scientific-pipelines-service-cli.
Running Tests Locally
Run ./gradlew service:test to run tests
Note: If you encounter errors indicating a failure to load the ApplicationContext due to an error while preparing a database cluster caused by a missing Docker environment,
this may be related to newer Docker versions (for example, 29.0.0 and above). To resolve this issue, override the
Docker API version in the $HOME/.docker-java.properties file. If the file does not already exist, create it and add the following line:
api.version=1.44
If the file mentioned already exists with above line, and the tests are still failing in the same way, try restarting Docker.
Running Linter Locally
Run ./gradlew spotlessCheck to run linter checks
Run ./gradlew :service:spotlessApply to apply fix any issues the linter finds
(Optional) Install pre-commit hooks
[scripts/git-hooks/pre-commit] has been provided to help ensure all submitted changes are formatted correctly. To install all hooks in [scripts/git-hooks], run:
git config core.hooksPath scripts/git-hooks
Running SonarQube locally
SonarQube is a static analysis code that scans code for a wide
range of issues, including maintainability and possible bugs. Get more information from
DSP SonarQube Docs
If you get a build failure due to
SonarQube and want to debug the problem locally, you need to get the sonar token from GSM
before running the gradle task.
export SONAR_TOKEN=$(gcloud secrets versions access latest --project="broad-dsde-dev" --secret="teaspoons-sonarcloud" | jq '.sonar_token')
./gradlew sonarqube
Running this task produces no output unless your project has errors. To
generate a report, run using --info:
./gradlew sonarqube --info
Connecting to the database
To connect to the Teaspoons database, we have a script in dsp-scripts that
does all the setup for you. Clone that repo and make sure you're either on Broad Internal wifi or connected
to the VPN. Then run the following command:
./db/psql-connect.sh dev teaspoons
Deploying to dev
Upon merging to main, the dev environment will be automatically deployed via the GitHub Action Bump, Tag, Publish, and Deploy
(that workflow is defined here).
The two tasks report-to-sherlock and set-version-in-dev will prompt Sherlock to deploy the new version to dev.
You can check the status of the deployment in Beehive and in
ArgoCD.
For more information about deployment to dev, check out DevOps' excellent documentation.
Tracing
We use OpenTelemetry for tracing, so that every request has a tracing span that can
be viewed in Google Cloud Trace.
See this DSP blog post for more info.
Running the BEE end-to-end tests
The end-to-end test that runs against a BEE is specified in .github/workflows/run-bee-e2e-tests.yaml. It calls the workflow defined
in the terra-github-workflows repo.
The end-to-end test is automatically run nightly on the dev environment.
To run the test against a specific feature branch:
Grab the image tag for your feature branch.
If you've opened a PR, you can find the image tag as follows:
go to the Bump, Tag, Publish, and Deploy workflow that's triggered each time you push to your branch
From there, go to the tag-publish-docker-deploy task
Expand the "Construct docker image name and tag" step
The first line should contain the image tag, something like "0.0.81-6761487".
Navigate to the e2e-test GHA workflow
Click on the "Run workflow" button and select your branch from the dropdown
Enter the image tag from step 1 in the "Custom image tag" field
If you've updated the end-to-end test in the dsp-resuable-workflows repo, enter either a commit hash or your git
branch name. If you don't need to change the test, leave the default as main.
Click the green "Run workflow" button.
Python clients
We publish a "thin", auto-generated Python client that wraps the Teaspoons APIs. This client is published to
PyPi and can be installed with
pip install teaspoons_client, although this is not meant to be user-facing. The thin api client is generated from
the OpenAPI spec in the openapi directory.
Publishing occurs automatically when a new version of the service is deployed, via the
release-python-client GHA.
We also have a user-facing, "thick" CLI whose code lives in a separate repository: DataBiosphere/terra-scientific-pipelines-service-cli. Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run them… Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run themselves in Terra. The most common reason for this is that the pipeline
accesses proprietary data that users are not allowed to access directly, but that may be used as e.g. a reference panel
for imputation.
Supported pipelines
Current supported pipelines are:
Array Imputation with the All of Us + AnVIL Reference Panel
Architecture
Architecture Doc
Architecture Diagram
Development
This codebase is in initial development.
Requirements
Technical
This service is written in Java 17, and uses Postgres 15.
To run locally, you'll also need:
jq - install with brew install jq
Java 17 - can be installed manually or through IntelliJ which will do it for you when importing the project
Postgres 15 - multiple solutions here as long as you have a postgres instance running on localhost:5432 the local app will connect appropriately. Be sure to use Postgres 15 (as of Feb 2025, Postgres 17 did not work)
Download Postgres.app (recommended) from https://postgresapp.com/
Brew https://formulae.brew.sh/formula/postgresql@15
External Services
Terra services
Sam
Used to authn users connecting to the service and authz users for admin endpoints
Rawls
Used to handle workspace interactions
creating methods
data tables
workflow submission
Cromwell
Used through Rawls to run submissions
Thurloe
Used to send notification emails to users
Tech stack
Java 17 temurin
Postgres 15
Gradle - build automation tool
SonarQube - static code security and coverage
Trivy - security scanner for docker images
Jib - docker image builder for Java
Local development
To run locally:
Make sure you have the requirements installed from above. We recommend IntelliJ as an IDE.
Clone the repo (if you see broken inputs build the project to get the generated sources)
Spin up a local postgres instance (NOTE: use version 15)
Run the commands in scripts/postgres-init.sql in your local postgres instance. You will need to be authenticated to access GSM.
Run scripts/write-config.sh
Run ./gradlew bootRun to spin up the server.
Navigate to http://localhost:8080/#
If this is your first time deploying to any environment, be sure to use the admin endpoint /api/admin/v1/pipelines/{pipelineName}/{pipelineVersion} to set your pipeline's workspace id.
To run this endpoint, you need to be authenticated using your firecloud test account. A list of accounts that developers typically need is here. Further, a list of resources that are generally useful is stored here
This endpoint requires two parameters directly, and three in the message body:
pipelineName can be retrieved by querying the /api/pipelines/v1 endpoint.
pipelineVersion can also be retrieved from the /api/pipelines/v1 endpoint.
workspaceBillingProject is listed in the Teaspoons Resources document linked above
workspaceName is also listed in the Teaspoons Resources document, and can be found through the Terra UI workspace dashboard
wdlMethodVersion is found for the specific workflow as listed in the Terra UI page for workflows.
Back up local Postgres databases before testing/refactors
Before running local migrations/refactors, take backups of both local databases so you can restore quickly.
Defaults in this repo (see service/src/main/resources/application.yml and scripts/postgres-init.sql):
host 127.0.0.1, port 5432
pipelines_db user/pass: dbuser / dbpwd
teaspoons_stairway_db user/pass: stairwayuser / stairwaypwd
Backup and verify:
ts="$(date +%Y%m%d_%H%M%S)"
backup_dir="$HOME/teaspoons-db-backups/$ts"
mkdir -p "$backup_dir"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_dump -Fc -f "$backup_dir/pipelines_db.dump" pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_dump -Fc -f "$backup_dir/teaspoons_stairway_db.dump" teaspoons_stairway_db
pg_restore -l "$backup_dir/pipelines_db.dump" | head
pg_restore -l "$backup_dir/teaspoons_stairway_db.dump" | head
echo "Backups written to: $backup_dir"
Restore later (replace <admin_password>):
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O dbuser pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_restore --clean --if-exists --no-owner -d pipelines_db "$backup_dir/pipelines_db.dump"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O stairwayuser teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_restore --clean --if-exists --no-owner -d teaspoons_stairway_db "$backup_dir/teaspoons_stairway_db.dump"
Local development with the UI
When running terra-ui locally against a local teaspoons backend, CORS-related errors can arise. To get around this, run the following command to copy a configuration file that allows requests from localhost:
./scripts/local-dev/copy_web_config.sh
Note that this file at the destination path (next to App.java) is ignored via .gitignore, since it should not be used in deployed environments.
Local development with debugging
If using Intellij (only IDE we use on the team), you can run the server with a debugger. Follow
the steps above but instead of running ./gradlew bootRun to spin up the server, you can run
(debug) the App.java class through intellij and set breakpoints in the code. Be sure to set the
GOOGLE_APPLICATION_CREDENTIALS=config/teaspoons-sa.json in the Run/Debug configuration Environment Variables.
Testing the CLI locally
If you make changes to openapi.yml, you should test the CLI locally.
To create the autogenerated Python client files locally, run
./gradlew :python-client:openApiGenerate
The files will be generated in python-client/generated and are ignored from being checked into the repo.
(Note: the unqualified ./gradlew openApiGenerate now regenerates all four codegen modules —
python-client, rawls-client, client, and service — so qualify the task when you only want the Python client.)
To test with the CLI, follow the instructions in the CLI repo: DataBiosphere/terra-scientific-pipelines-service-cli.
Running Tests Locally
Run ./gradlew service:test to run tests
Note: If you encounter errors indicating a failure to load the ApplicationContext due to an error while preparing a database cluster caused by a missing Docker environment,
this may be related to newer Docker versions (for example, 29.0.0 and above). To resolve this issue, override the
Docker API version in the $HOME/.docker-java.properties file. If the file does not already exist, create it and add the following line:
api.version=1.44
If the file mentioned already exists with above line, and the tests are still failing in the same way, try restarting Docker.
Running Linter Locally
Run ./gradlew spotlessCheck to run linter checks
Run ./gradlew :service:spotlessApply to apply fix any issues the linter finds
(Optional) Install pre-commit hooks
[scripts/git-hooks/pre-commit] has been provided to help ensure all submitted changes are formatted correctly. To install all hooks in [scripts/git-hooks], run:
git config core.hooksPath scripts/git-hooks
Running SonarQube locally
SonarQube is a static analysis code that scans code for a wide
range of issues, including maintainability and possible bugs. Get more information from
DSP SonarQube Docs
If you get a build failure due to
SonarQube and want to debug the problem locally, you need to get the sonar token from GSM
before running the gradle task.
export SONAR_TOKEN=$(gcloud secrets versions access latest --project="broad-dsde-dev" --secret="teaspoons-sonarcloud" | jq '.sonar_token')
./gradlew sonarqube
Running this task produces no output unless your project has errors. To
generate a report, run using --info:
./gradlew sonarqube --info
Connecting to the database
To connect to the Teaspoons database, we have a script in dsp-scripts that
does all the setup for you. Clone that repo and make sure you're either on Broad Internal wifi or connected
to the VPN. Then run the following command:
./db/psql-connect.sh dev teaspoons
Deploying to dev
Upon merging to main, the dev environment will be automatically deployed via the GitHub Action Bump, Tag, Publish, and Deploy
(that workflow is defined here).
The two tasks report-to-sherlock and set-version-in-dev will prompt Sherlock to deploy the new version to dev.
You can check the status of the deployment in Beehive and in
ArgoCD.
For more information about deployment to dev, check out DevOps' excellent documentation.
Tracing
We use OpenTelemetry for tracing, so that every request has a tracing span that can
be viewed in Google Cloud Trace.
See this DSP blog post for more info.
Running the BEE end-to-end tests
The end-to-end test that runs against a BEE is specified in .github/workflows/run-bee-e2e-tests.yaml. It calls the workflow defined
in the terra-github-workflows repo.
The end-to-end test is automatically run nightly on the dev environment.
To run the test against a specific feature branch:
Grab the image tag for your feature branch.
If you've opened a PR, you can find the image tag as follows:
go to the Bump, Tag, Publish, and Deploy workflow that's triggered each time you push to your branch
From there, go to the tag-publish-docker-deploy task
Expand the "Construct docker image name and tag" step
The first line should contain the image tag, something like "0.0.81-6761487".
Navigate to the e2e-test GHA workflow
Click on the "Run workflow" button and select your branch from the dropdown
Enter the image tag from step 1 in the "Custom image tag" field
If you've updated the end-to-end test in the dsp-resuable-workflows repo, enter either a commit hash or your git
branch name. If you don't need to change the test, leave the default as main.
Click the green "Run workflow" button.
Python clients
We publish a "thin", auto-generated Python client that wraps the Teaspoons APIs. This client is published to
PyPi and can be installed with
pip install teaspoons_client, although this is not meant to be user-facing. The thin api client is generated from
the OpenAPI spec in the openapi directory.
Publishing occurs automatically when a new version of the service is deployed, via the
release-python-client GHA.
We also have a user-facing, "thick" CLI whose code lives in a separate repository: DataBiosphere/terra-scientific-pipelines-service-cli. Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run them… Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run themselves in Terra. The most common reason for this is that the pipeline
accesses proprietary data that users are not allowed to access directly, but that may be used as e.g. a reference panel
for imputation.
Supported pipelines
Current supported pipelines are:
Array Imputation with the All of Us + AnVIL Reference Panel
Architecture
Architecture Doc
Architecture Diagram
Development
This codebase is in initial development.
Requirements
Technical
This service is written in Java 17, and uses Postgres 15.
To run locally, you'll also need:
jq - install with brew install jq
Java 17 - can be installed manually or through IntelliJ which will do it for you when importing the project
Postgres 15 - multiple solutions here as long as you have a postgres instance running on localhost:5432 the local app will connect appropriately. Be sure to use Postgres 15 (as of Feb 2025, Postgres 17 did not work)
Download Postgres.app (recommended) from https://postgresapp.com/
Brew https://formulae.brew.sh/formula/postgresql@15
External Services
Terra services
Sam
Used to authn users connecting to the service and authz users for admin endpoints
Rawls
Used to handle workspace interactions
creating methods
data tables
workflow submission
Cromwell
Used through Rawls to run submissions
Thurloe
Used to send notification emails to users
Tech stack
Java 17 temurin
Postgres 15
Gradle - build automation tool
SonarQube - static code security and coverage
Trivy - security scanner for docker images
Jib - docker image builder for Java
Local development
To run locally:
Make sure you have the requirements installed from above. We recommend IntelliJ as an IDE.
Clone the repo (if you see broken inputs build the project to get the generated sources)
Spin up a local postgres instance (NOTE: use version 15)
Run the commands in scripts/postgres-init.sql in your local postgres instance. You will need to be authenticated to access GSM.
Run scripts/write-config.sh
Run ./gradlew bootRun to spin up the server.
Navigate to http://localhost:8080/#
If this is your first time deploying to any environment, be sure to use the admin endpoint /api/admin/v1/pipelines/{pipelineName}/{pipelineVersion} to set your pipeline's workspace id.
To run this endpoint, you need to be authenticated using your firecloud test account. A list of accounts that developers typically need is here. Further, a list of resources that are generally useful is stored here
This endpoint requires two parameters directly, and three in the message body:
pipelineName can be retrieved by querying the /api/pipelines/v1 endpoint.
pipelineVersion can also be retrieved from the /api/pipelines/v1 endpoint.
workspaceBillingProject is listed in the Teaspoons Resources document linked above
workspaceName is also listed in the Teaspoons Resources document, and can be found through the Terra UI workspace dashboard
wdlMethodVersion is found for the specific workflow as listed in the Terra UI page for workflows.
Back up local Postgres databases before testing/refactors
Before running local migrations/refactors, take backups of both local databases so you can restore quickly.
Defaults in this repo (see service/src/main/resources/application.yml and scripts/postgres-init.sql):
host 127.0.0.1, port 5432
pipelines_db user/pass: dbuser / dbpwd
teaspoons_stairway_db user/pass: stairwayuser / stairwaypwd
Backup and verify:
ts="$(date +%Y%m%d_%H%M%S)"
backup_dir="$HOME/teaspoons-db-backups/$ts"
mkdir -p "$backup_dir"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_dump -Fc -f "$backup_dir/pipelines_db.dump" pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_dump -Fc -f "$backup_dir/teaspoons_stairway_db.dump" teaspoons_stairway_db
pg_restore -l "$backup_dir/pipelines_db.dump" | head
pg_restore -l "$backup_dir/teaspoons_stairway_db.dump" | head
echo "Backups written to: $backup_dir"
Restore later (replace <admin_password>):
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O dbuser pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_restore --clean --if-exists --no-owner -d pipelines_db "$backup_dir/pipelines_db.dump"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O stairwayuser teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_restore --clean --if-exists --no-owner -d teaspoons_stairway_db "$backup_dir/teaspoons_stairway_db.dump"
Local development with the UI
When running terra-ui locally against a local teaspoons backend, CORS-related errors can arise. To get around this, run the following command to copy a configuration file that allows requests from localhost:
./scripts/local-dev/copy_web_config.sh
Note that this file at the destination path (next to App.java) is ignored via .gitignore, since it should not be used in deployed environments.
Local development with debugging
If using Intellij (only IDE we use on the team), you can run the server with a debugger. Follow
the steps above but instead of running ./gradlew bootRun to spin up the server, you can run
(debug) the App.java class through intellij and set breakpoints in the code. Be sure to set the
GOOGLE_APPLICATION_CREDENTIALS=config/teaspoons-sa.json in the Run/Debug configuration Environment Variables.
Testing the CLI locally
If you make changes to openapi.yml, you should test the CLI locally.
To create the autogenerated Python client files locally, run
./gradlew :python-client:openApiGenerate
The files will be generated in python-client/generated and are ignored from being checked into the repo.
(Note: the unqualified ./gradlew openApiGenerate now regenerates all four codegen modules —
python-client, rawls-client, client, and service — so qualify the task when you only want the Python client.)
To test with the CLI, follow the instructions in the CLI repo: DataBiosphere/terra-scientific-pipelines-service-cli.
Running Tests Locally
Run ./gradlew service:test to run tests
Note: If you encounter errors indicating a failure to load the ApplicationContext due to an error while preparing a database cluster caused by a missing Docker environment,
this may be related to newer Docker versions (for example, 29.0.0 and above). To resolve this issue, override the
Docker API version in the $HOME/.docker-java.properties file. If the file does not already exist, create it and add the following line:
api.version=1.44
If the file mentioned already exists with above line, and the tests are still failing in the same way, try restarting Docker.
Running Linter Locally
Run ./gradlew spotlessCheck to run linter checks
Run ./gradlew :service:spotlessApply to apply fix any issues the linter finds
(Optional) Install pre-commit hooks
[scripts/git-hooks/pre-commit] has been provided to help ensure all submitted changes are formatted correctly. To install all hooks in [scripts/git-hooks], run:
git config core.hooksPath scripts/git-hooks
Running SonarQube locally
SonarQube is a static analysis code that scans code for a wide
range of issues, including maintainability and possible bugs. Get more information from
DSP SonarQube Docs
If you get a build failure due to
SonarQube and want to debug the problem locally, you need to get the sonar token from GSM
before running the gradle task.
export SONAR_TOKEN=$(gcloud secrets versions access latest --project="broad-dsde-dev" --secret="teaspoons-sonarcloud" | jq '.sonar_token')
./gradlew sonarqube
Running this task produces no output unless your project has errors. To
generate a report, run using --info:
./gradlew sonarqube --info
Connecting to the database
To connect to the Teaspoons database, we have a script in dsp-scripts that
does all the setup for you. Clone that repo and make sure you're either on Broad Internal wifi or connected
to the VPN. Then run the following command:
./db/psql-connect.sh dev teaspoons
Deploying to dev
Upon merging to main, the dev environment will be automatically deployed via the GitHub Action Bump, Tag, Publish, and Deploy
(that workflow is defined here).
The two tasks report-to-sherlock and set-version-in-dev will prompt Sherlock to deploy the new version to dev.
You can check the status of the deployment in Beehive and in
ArgoCD.
For more information about deployment to dev, check out DevOps' excellent documentation.
Tracing
We use OpenTelemetry for tracing, so that every request has a tracing span that can
be viewed in Google Cloud Trace.
See this DSP blog post for more info.
Running the BEE end-to-end tests
The end-to-end test that runs against a BEE is specified in .github/workflows/run-bee-e2e-tests.yaml. It calls the workflow defined
in the terra-github-workflows repo.
The end-to-end test is automatically run nightly on the dev environment.
To run the test against a specific feature branch:
Grab the image tag for your feature branch.
If you've opened a PR, you can find the image tag as follows:
go to the Bump, Tag, Publish, and Deploy workflow that's triggered each time you push to your branch
From there, go to the tag-publish-docker-deploy task
Expand the "Construct docker image name and tag" step
The first line should contain the image tag, something like "0.0.81-6761487".
Navigate to the e2e-test GHA workflow
Click on the "Run workflow" button and select your branch from the dropdown
Enter the image tag from step 1 in the "Custom image tag" field
If you've updated the end-to-end test in the dsp-resuable-workflows repo, enter either a commit hash or your git
branch name. If you don't need to change the test, leave the default as main.
Click the green "Run workflow" button.
Python clients
We publish a "thin", auto-generated Python client that wraps the Teaspoons APIs. This client is published to
PyPi and can be installed with
pip install teaspoons_client, although this is not meant to be user-facing. The thin api client is generated from
the OpenAPI spec in the openapi directory.
Publishing occurs automatically when a new version of the service is deployed, via the
release-python-client GHA.
We also have a user-facing, "thick" CLI whose code lives in a separate repository: DataBiosphere/terra-scientific-pipelines-service-cli. 🌙 AMIGOpy Pre-Release / Development Build
Release Tag: nightly_28092026
Commit: fafdde965fd8fa7aeb9adf47e59bcea9f20a971c
Status: ⚠️ Bleeding-edge pre-release build (Unsigned).
📥 Downloads
Installer… 🌙 AMIGOpy Pre-Release / Development Build
Release Tag: nightly_28092026
Commit: fafdde965fd8fa7aeb9adf47e59bcea9f20a971c
Status: ⚠️ Bleeding-edge pre-release build (Unsigned).
📥 Downloads
Installer: AMIGOpy_Setup_nightly_28092026.exe
Portable ZIP: AMIGOpy-win-portable-nightly_28092026.zip Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run them… Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run themselves in Terra. The most common reason for this is that the pipeline
accesses proprietary data that users are not allowed to access directly, but that may be used as e.g. a reference panel
for imputation.
Supported pipelines
Current supported pipelines are:
Array Imputation with the All of Us + AnVIL Reference Panel
Architecture
Architecture Doc
Architecture Diagram
Development
This codebase is in initial development.
Requirements
Technical
This service is written in Java 17, and uses Postgres 15.
To run locally, you'll also need:
jq - install with brew install jq
Java 17 - can be installed manually or through IntelliJ which will do it for you when importing the project
Postgres 15 - multiple solutions here as long as you have a postgres instance running on localhost:5432 the local app will connect appropriately. Be sure to use Postgres 15 (as of Feb 2025, Postgres 17 did not work)
Download Postgres.app (recommended) from https://postgresapp.com/
Brew https://formulae.brew.sh/formula/postgresql@15
External Services
Terra services
Sam
Used to authn users connecting to the service and authz users for admin endpoints
Rawls
Used to handle workspace interactions
creating methods
data tables
workflow submission
Cromwell
Used through Rawls to run submissions
Thurloe
Used to send notification emails to users
Tech stack
Java 17 temurin
Postgres 15
Gradle - build automation tool
SonarQube - static code security and coverage
Trivy - security scanner for docker images
Jib - docker image builder for Java
Local development
To run locally:
Make sure you have the requirements installed from above. We recommend IntelliJ as an IDE.
Clone the repo (if you see broken inputs build the project to get the generated sources)
Spin up a local postgres instance (NOTE: use version 15)
Run the commands in scripts/postgres-init.sql in your local postgres instance. You will need to be authenticated to access GSM.
Run scripts/write-config.sh
Run ./gradlew bootRun to spin up the server.
Navigate to http://localhost:8080/#
If this is your first time deploying to any environment, be sure to use the admin endpoint /api/admin/v1/pipelines/{pipelineName}/{pipelineVersion} to set your pipeline's workspace id.
To run this endpoint, you need to be authenticated using your firecloud test account. A list of accounts that developers typically need is here. Further, a list of resources that are generally useful is stored here
This endpoint requires two parameters directly, and three in the message body:
pipelineName can be retrieved by querying the /api/pipelines/v1 endpoint.
pipelineVersion can also be retrieved from the /api/pipelines/v1 endpoint.
workspaceBillingProject is listed in the Teaspoons Resources document linked above
workspaceName is also listed in the Teaspoons Resources document, and can be found through the Terra UI workspace dashboard
wdlMethodVersion is found for the specific workflow as listed in the Terra UI page for workflows.
Back up local Postgres databases before testing/refactors
Before running local migrations/refactors, take backups of both local databases so you can restore quickly.
Defaults in this repo (see service/src/main/resources/application.yml and scripts/postgres-init.sql):
host 127.0.0.1, port 5432
pipelines_db user/pass: dbuser / dbpwd
teaspoons_stairway_db user/pass: stairwayuser / stairwaypwd
Backup and verify:
ts="$(date +%Y%m%d_%H%M%S)"
backup_dir="$HOME/teaspoons-db-backups/$ts"
mkdir -p "$backup_dir"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_dump -Fc -f "$backup_dir/pipelines_db.dump" pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_dump -Fc -f "$backup_dir/teaspoons_stairway_db.dump" teaspoons_stairway_db
pg_restore -l "$backup_dir/pipelines_db.dump" | head
pg_restore -l "$backup_dir/teaspoons_stairway_db.dump" | head
echo "Backups written to: $backup_dir"
Restore later (replace <admin_password>):
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O dbuser pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_restore --clean --if-exists --no-owner -d pipelines_db "$backup_dir/pipelines_db.dump"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O stairwayuser teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_restore --clean --if-exists --no-owner -d teaspoons_stairway_db "$backup_dir/teaspoons_stairway_db.dump"
Local development with the UI
When running terra-ui locally against a local teaspoons backend, CORS-related errors can arise. To get around this, run the following command to copy a configuration file that allows requests from localhost:
./scripts/local-dev/copy_web_config.sh
Note that this file at the destination path (next to App.java) is ignored via .gitignore, since it should not be used in deployed environments.
Local development with debugging
If using Intellij (only IDE we use on the team), you can run the server with a debugger. Follow
the steps above but instead of running ./gradlew bootRun to spin up the server, you can run
(debug) the App.java class through intellij and set breakpoints in the code. Be sure to set the
GOOGLE_APPLICATION_CREDENTIALS=config/teaspoons-sa.json in the Run/Debug configuration Environment Variables.
Testing the CLI locally
If you make changes to openapi.yml, you should test the CLI locally.
To create the autogenerated Python client files locally, run
./gradlew :python-client:openApiGenerate
The files will be generated in python-client/generated and are ignored from being checked into the repo.
(Note: the unqualified ./gradlew openApiGenerate now regenerates all four codegen modules —
python-client, rawls-client, client, and service — so qualify the task when you only want the Python client.)
To test with the CLI, follow the instructions in the CLI repo: DataBiosphere/terra-scientific-pipelines-service-cli.
Running Tests Locally
Run ./gradlew service:test to run tests
Note: If you encounter errors indicating a failure to load the ApplicationContext due to an error while preparing a database cluster caused by a missing Docker environment,
this may be related to newer Docker versions (for example, 29.0.0 and above). To resolve this issue, override the
Docker API version in the $HOME/.docker-java.properties file. If the file does not already exist, create it and add the following line:
api.version=1.44
If the file mentioned already exists with above line, and the tests are still failing in the same way, try restarting Docker.
Running Linter Locally
Run ./gradlew spotlessCheck to run linter checks
Run ./gradlew :service:spotlessApply to apply fix any issues the linter finds
(Optional) Install pre-commit hooks
[scripts/git-hooks/pre-commit] has been provided to help ensure all submitted changes are formatted correctly. To install all hooks in [scripts/git-hooks], run:
git config core.hooksPath scripts/git-hooks
Running SonarQube locally
SonarQube is a static analysis code that scans code for a wide
range of issues, including maintainability and possible bugs. Get more information from
DSP SonarQube Docs
If you get a build failure due to
SonarQube and want to debug the problem locally, you need to get the sonar token from GSM
before running the gradle task.
export SONAR_TOKEN=$(gcloud secrets versions access latest --project="broad-dsde-dev" --secret="teaspoons-sonarcloud" | jq '.sonar_token')
./gradlew sonarqube
Running this task produces no output unless your project has errors. To
generate a report, run using --info:
./gradlew sonarqube --info
Connecting to the database
To connect to the Teaspoons database, we have a script in dsp-scripts that
does all the setup for you. Clone that repo and make sure you're either on Broad Internal wifi or connected
to the VPN. Then run the following command:
./db/psql-connect.sh dev teaspoons
Deploying to dev
Upon merging to main, the dev environment will be automatically deployed via the GitHub Action Bump, Tag, Publish, and Deploy
(that workflow is defined here).
The two tasks report-to-sherlock and set-version-in-dev will prompt Sherlock to deploy the new version to dev.
You can check the status of the deployment in Beehive and in
ArgoCD.
For more information about deployment to dev, check out DevOps' excellent documentation.
Tracing
We use OpenTelemetry for tracing, so that every request has a tracing span that can
be viewed in Google Cloud Trace.
See this DSP blog post for more info.
Running the BEE end-to-end tests
The end-to-end test that runs against a BEE is specified in .github/workflows/run-bee-e2e-tests.yaml. It calls the workflow defined
in the terra-github-workflows repo.
The end-to-end test is automatically run nightly on the dev environment.
To run the test against a specific feature branch:
Grab the image tag for your feature branch.
If you've opened a PR, you can find the image tag as follows:
go to the Bump, Tag, Publish, and Deploy workflow that's triggered each time you push to your branch
From there, go to the tag-publish-docker-deploy task
Expand the "Construct docker image name and tag" step
The first line should contain the image tag, something like "0.0.81-6761487".
Navigate to the e2e-test GHA workflow
Click on the "Run workflow" button and select your branch from the dropdown
Enter the image tag from step 1 in the "Custom image tag" field
If you've updated the end-to-end test in the dsp-resuable-workflows repo, enter either a commit hash or your git
branch name. If you don't need to change the test, leave the default as main.
Click the green "Run workflow" button.
Python clients
We publish a "thin", auto-generated Python client that wraps the Teaspoons APIs. This client is published to
PyPi and can be installed with
pip install teaspoons_client, although this is not meant to be user-facing. The thin api client is generated from
the OpenAPI spec in the openapi directory.
Publishing occurs automatically when a new version of the service is deployed, via the
release-python-client GHA.
We also have a user-facing, "thick" CLI whose code lives in a separate repository: DataBiosphere/terra-scientific-pipelines-service-cli. Ushbu maqolada mahalliy o‘simlik va agro-sanoat xomashyolari asosida ekologik xavfsiz korroziya ingibitorlari ishlab chiqarishning zamonaviy yondashuvlari tahlil qilindi. Tadqiqot ad… Ushbu maqolada mahalliy o‘simlik va agro-sanoat xomashyolari asosida ekologik xavfsiz korroziya ingibitorlari ishlab chiqarishning zamonaviy yondashuvlari tahlil qilindi. Tadqiqot adabiyotlar sharhi va analitik sintez shaklida bajarilib, uzum urug‘i va po‘stlog‘i, anor po‘stlog‘i, yong‘oq yashil po‘stlog‘i, tannin va ligninga boy qishloq xo‘jaligi chiqindilarining uglerodli va kam legirlangan po‘latlar korroziyasini kamaytirishdagi imkoniyatlari baholandi. Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run them… Terra Scientific Pipelines Service
Overview
Terra Scientific Pipelines Service, or Teaspoons, facilitates running a number of defined scientific pipelines
on behalf of users that users can't run themselves in Terra. The most common reason for this is that the pipeline
accesses proprietary data that users are not allowed to access directly, but that may be used as e.g. a reference panel
for imputation.
Supported pipelines
Current supported pipelines are:
Array Imputation with the All of Us + AnVIL Reference Panel
Architecture
Architecture Doc
Architecture Diagram
Development
This codebase is in initial development.
Requirements
Technical
This service is written in Java 17, and uses Postgres 15.
To run locally, you'll also need:
jq - install with brew install jq
Java 17 - can be installed manually or through IntelliJ which will do it for you when importing the project
Postgres 15 - multiple solutions here as long as you have a postgres instance running on localhost:5432 the local app will connect appropriately. Be sure to use Postgres 15 (as of Feb 2025, Postgres 17 did not work)
Download Postgres.app (recommended) from https://postgresapp.com/
Brew https://formulae.brew.sh/formula/postgresql@15
External Services
Terra services
Sam
Used to authn users connecting to the service and authz users for admin endpoints
Rawls
Used to handle workspace interactions
creating methods
data tables
workflow submission
Cromwell
Used through Rawls to run submissions
Thurloe
Used to send notification emails to users
Tech stack
Java 17 temurin
Postgres 15
Gradle - build automation tool
SonarQube - static code security and coverage
Trivy - security scanner for docker images
Jib - docker image builder for Java
Local development
To run locally:
Make sure you have the requirements installed from above. We recommend IntelliJ as an IDE.
Clone the repo (if you see broken inputs build the project to get the generated sources)
Spin up a local postgres instance (NOTE: use version 15)
Run the commands in scripts/postgres-init.sql in your local postgres instance. You will need to be authenticated to access GSM.
Run scripts/write-config.sh
Run ./gradlew bootRun to spin up the server.
Navigate to http://localhost:8080/#
If this is your first time deploying to any environment, be sure to use the admin endpoint /api/admin/v1/pipelines/{pipelineName}/{pipelineVersion} to set your pipeline's workspace id.
To run this endpoint, you need to be authenticated using your firecloud test account. A list of accounts that developers typically need is here. Further, a list of resources that are generally useful is stored here
This endpoint requires two parameters directly, and three in the message body:
pipelineName can be retrieved by querying the /api/pipelines/v1 endpoint.
pipelineVersion can also be retrieved from the /api/pipelines/v1 endpoint.
workspaceBillingProject is listed in the Teaspoons Resources document linked above
workspaceName is also listed in the Teaspoons Resources document, and can be found through the Terra UI workspace dashboard
wdlMethodVersion is found for the specific workflow as listed in the Terra UI page for workflows.
Back up local Postgres databases before testing/refactors
Before running local migrations/refactors, take backups of both local databases so you can restore quickly.
Defaults in this repo (see service/src/main/resources/application.yml and scripts/postgres-init.sql):
host 127.0.0.1, port 5432
pipelines_db user/pass: dbuser / dbpwd
teaspoons_stairway_db user/pass: stairwayuser / stairwaypwd
Backup and verify:
ts="$(date +%Y%m%d_%H%M%S)"
backup_dir="$HOME/teaspoons-db-backups/$ts"
mkdir -p "$backup_dir"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_dump -Fc -f "$backup_dir/pipelines_db.dump" pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_dump -Fc -f "$backup_dir/teaspoons_stairway_db.dump" teaspoons_stairway_db
pg_restore -l "$backup_dir/pipelines_db.dump" | head
pg_restore -l "$backup_dir/teaspoons_stairway_db.dump" | head
echo "Backups written to: $backup_dir"
Restore later (replace <admin_password>):
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O dbuser pipelines_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=dbuser PGPASSWORD=dbpwd \
pg_restore --clean --if-exists --no-owner -d pipelines_db "$backup_dir/pipelines_db.dump"
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> dropdb --if-exists teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=postgres PGPASSWORD=<admin_password> createdb -O stairwayuser teaspoons_stairway_db
PGHOST=127.0.0.1 PGPORT=5432 PGUSER=stairwayuser PGPASSWORD=stairwaypwd \
pg_restore --clean --if-exists --no-owner -d teaspoons_stairway_db "$backup_dir/teaspoons_stairway_db.dump"
Local development with the UI
When running terra-ui locally against a local teaspoons backend, CORS-related errors can arise. To get around this, run the following command to copy a configuration file that allows requests from localhost:
./scripts/local-dev/copy_web_config.sh
Note that this file at the destination path (next to App.java) is ignored via .gitignore, since it should not be used in deployed environments.
Local development with debugging
If using Intellij (only IDE we use on the team), you can run the server with a debugger. Follow
the steps above but instead of running ./gradlew bootRun to spin up the server, you can run
(debug) the App.java class through intellij and set breakpoints in the code. Be sure to set the
GOOGLE_APPLICATION_CREDENTIALS=config/teaspoons-sa.json in the Run/Debug configuration Environment Variables.
Testing the CLI locally
If you make changes to openapi.yml, you should test the CLI locally.
To create the autogenerated Python client files locally, run
./gradlew :python-client:openApiGenerate
The files will be generated in python-client/generated and are ignored from being checked into the repo.
(Note: the unqualified ./gradlew openApiGenerate now regenerates all four codegen modules —
python-client, rawls-client, client, and service — so qualify the task when you only want the Python client.)
To test with the CLI, follow the instructions in the CLI repo: DataBiosphere/terra-scientific-pipelines-service-cli.
Running Tests Locally
Run ./gradlew service:test to run tests
Note: If you encounter errors indicating a failure to load the ApplicationContext due to an error while preparing a database cluster caused by a missing Docker environment,
this may be related to newer Docker versions (for example, 29.0.0 and above). To resolve this issue, override the
Docker API version in the $HOME/.docker-java.properties file. If the file does not already exist, create it and add the following line:
api.version=1.44
If the file mentioned already exists with above line, and the tests are still failing in the same way, try restarting Docker.
Running Linter Locally
Run ./gradlew spotlessCheck to run linter checks
Run ./gradlew :service:spotlessApply to apply fix any issues the linter finds
(Optional) Install pre-commit hooks
[scripts/git-hooks/pre-commit] has been provided to help ensure all submitted changes are formatted correctly. To install all hooks in [scripts/git-hooks], run:
git config core.hooksPath scripts/git-hooks
Running SonarQube locally
SonarQube is a static analysis code that scans code for a wide
range of issues, including maintainability and possible bugs. Get more information from
DSP SonarQube Docs
If you get a build failure due to
SonarQube and want to debug the problem locally, you need to get the sonar token from GSM
before running the gradle task.
export SONAR_TOKEN=$(gcloud secrets versions access latest --project="broad-dsde-dev" --secret="teaspoons-sonarcloud" | jq '.sonar_token')
./gradlew sonarqube
Running this task produces no output unless your project has errors. To
generate a report, run using --info:
./gradlew sonarqube --info
Connecting to the database
To connect to the Teaspoons database, we have a script in dsp-scripts that
does all the setup for you. Clone that repo and make sure you're either on Broad Internal wifi or connected
to the VPN. Then run the following command:
./db/psql-connect.sh dev teaspoons
Deploying to dev
Upon merging to main, the dev environment will be automatically deployed via the GitHub Action Bump, Tag, Publish, and Deploy
(that workflow is defined here).
The two tasks report-to-sherlock and set-version-in-dev will prompt Sherlock to deploy the new version to dev.
You can check the status of the deployment in Beehive and in
ArgoCD.
For more information about deployment to dev, check out DevOps' excellent documentation.
Tracing
We use OpenTelemetry for tracing, so that every request has a tracing span that can
be viewed in Google Cloud Trace.
See this DSP blog post for more info.
Running the BEE end-to-end tests
The end-to-end test that runs against a BEE is specified in .github/workflows/run-bee-e2e-tests.yaml. It calls the workflow defined
in the terra-github-workflows repo.
The end-to-end test is automatically run nightly on the dev environment.
To run the test against a specific feature branch:
Grab the image tag for your feature branch.
If you've opened a PR, you can find the image tag as follows:
go to the Bump, Tag, Publish, and Deploy workflow that's triggered each time you push to your branch
From there, go to the tag-publish-docker-deploy task
Expand the "Construct docker image name and tag" step
The first line should contain the image tag, something like "0.0.81-6761487".
Navigate to the e2e-test GHA workflow
Click on the "Run workflow" button and select your branch from the dropdown
Enter the image tag from step 1 in the "Custom image tag" field
If you've updated the end-to-end test in the dsp-resuable-workflows repo, enter either a commit hash or your git
branch name. If you don't need to change the test, leave the default as main.
Click the green "Run workflow" button.
Python clients
We publish a "thin", auto-generated Python client that wraps the Teaspoons APIs. This client is published to
PyPi and can be installed with
pip install teaspoons_client, although this is not meant to be user-facing. The thin api client is generated from
the OpenAPI spec in the openapi directory.
Publishing occurs automatically when a new version of the service is deployed, via the
release-python-client GHA.
We also have a user-facing, "thick" CLI whose code lives in a separate repository: DataBiosphere/terra-scientific-pipelines-service-cli.github.com/DataBiosphere/terra-scientific-pipelines-service/RelocateAllSVReferencePanelFiles
GalaxyMemMD: Galaxy workflows, MD trajectories, and starting structures for reproducible membrane protein embedding, simulation, and analysis
github.com/DataBiosphere/terra-scientific-pipelines-service/RecombineVariantAndHomRefVcfs
github.com/DataBiosphere/terra-scientific-pipelines-service/QuotaConsumedEmpty
A Class of Magnetic Topological Material Candidates with Hypervalent Bi Chains
github.com/DataBiosphere/terra-scientific-pipelines-service/MakeSitesOnly
PhysicsResearch/AMIGOpy: AMIGOpy Pre-Release (nightly_28092026)
github.com/DataBiosphere/terra-scientific-pipelines-service/LiftoverVcfs
MAHALLIY XOMASHYOLAR ASOSIDA KORROZIYAGA QARSHI MODDALAR ISHLAB CHIQARISH
github.com/DataBiosphere/terra-scientific-pipelines-service/LeftAlignVcf
On Losses, Pauses, Jumps and the Wideband E-Model – IEEE Xplore Document
There is an increasing interest in upgrading the EModel, a parametric tool for speech quality estimation, to the wideband and super-wideband contexts. The
NUAV – a testbed for developing autonomous Unmanned Aerial Vehicles – IEEE Xplore Document
Contemporary models of Unmanned Aerial Vehicles (UAVs) are largely developed using simulators. In a typical scheme, a flight simulator is dovetailed with a
NUAV – a testbed for developing autonomous Unmanned Aerial Vehicles
Simulators as Drivers of Cutting Edge Research – IEEE Xplore Document
Undertaking engineering research can be compounding for beginning graduate students and thwarting even for seasoned researchers. With a wealth of academic
Simulators as Drivers of Cutting Edge Research
Evolutionary speech quality estimation in VoIP
A Methodology for Deriving VoIP Equipment Impairment Factors for a Mixed NB/WB Context
Real-Time, Non-intrusive Speech Quality Estimation: A Signal-Based Mod
Real-Time, Non-intrusive Evaluation of VoIP
VoIP speech quality estimation in a mixed context with genetic programming
An Evolutionary Approach to Speech Quality Estimation
Real-Time Non-Intrusive VoIP Evaluation Using Second Generation Network Processor
Non-intrusive quality evaluation of VoIP using genetic programming
