GenotypeGVCFs task consistently fails with VMReportingTimeout (50002) / exit code 247 on GCP Batch — joint genotyping ~400 exomes

Post author
Akili Joseph

I am running a joint-genotyping workflow (GenomicsDBImport → GenotypeGVCFs → hard-filter → gather) on ~400 whole-exome gVCFs in a Terra workspace on the GCP Batch backend. GATK 4.5.0.0.

GenomicsDBImport succeeds for all interval shards however GenotypeGVCFs consistently fails, across multiple configurations, with:

  •  `Job exit code 247`, and on several shards:
  • `The job was stopped before the command finished. GCP Batch task exited with VMReportingTimeout(50002).`


Reading a failed shard's stderr, GenotypeGVCFs is running normally (ProgressMeter advancing, records being written e.g. reached chr11 after ~38 minutes, ~148,000 records) and then the log ends abruptly with no Java error, no OutOfMemoryError, and no exception. The output VCF is not produced. This is consistent with the VM being reaped by Batch rather than the task crashing.

What I have already tried (none resolved it):

  • Increased GenotypeGVCFs task memory from 8 GB → 26 GB → 52 GB (it never showed a Java OOM).
  • Increased scatter count from 24 → 50 → 150 intervals (so each task is small; still fails, still reaches only a low position before the VM is reaped).
  • Switched the task to non-preemptible.
  • Increased disk and moved to SSD (local-disk 100 SSD).
  • Reduced logging verbosity (`--verbosity WARNING`) to stop a heavy GenomicsDB per-iteration timer log flood I observed in stderr.


The failure persists regardless. GenomicsDBImport (same VMs, same data) succeeds, so it appears specific to GenotypeGVCFs reading the GenomicsDB, or to a transient Batch VM-health issue.

Questions:
1. Is this a known transient GCP Batch issue for this workspace's region/billing project, and is
   there a recommended region or configuration change?
2. What Cromwell version is this workspace on? I see that Cromwell 89+ added automatic transient retries for Batch 50002 errors and Cromwell 91 improved this further. Are those retries active for mid-run 50002 failures in my workspace, and if not, how do I enable them?
3. Is there a recommended `maxRetries` / workflow-options configuration to make tasks that hit 50002 retry automatically rather than failing the workflow?
4. Are there logs (report agent state entries) I should extract to help diagnose whether the VM is hanging on memory pressure vs. an I/O stall on the GenomicsDB read?

Any guidance on resolving the 50002 / making GenotypeGVCFs complete on ~400 exomes would be greatly appreciated. Thank you.

Comments

0 comments

Please sign in to leave a comment.