Executive Summary
A US card issuer still posted, settled, and extracted general-ledger and BSA files from a z/OS batch estate: 3,800 COBOL programs, 890 copybooks, 1,420 JCL members, VSAM plus DB2, 6.2 million lines counted once. This was not CICS online and it was not a rehost. The bank already ran Databricks for fraud features. Nightly card-posting and collections batch was the thing still buying MIPS. MigryX parsed copybooks into Spark schemas (COMP-3, REDEFINES, OCCURS), rewrote PROCEDURE DIVISION as PySpark DataFrames — not record loops — and turned JCL EXEC/DD/PROC graphs into Databricks Workflows. Twenty-two months. Dual-run through two quarter-ends. The LPAR that existed for this batch is gone. Three-year gap versus MIPS + OEM tools: $11.4 million.
Client Overview
Issuing, collections, and finance ops grew up on the mainframe because that is where the card system of record lived. Authorization is a different stack. This program was the overnight work: post presentments, age delinquencies, cut GL, feed BSA, produce investor tapes. Analysts who joined after 2018 expected Python. The people who could change a copybook were retiring on a calendar, not a slogan.
Databricks was already paid for. Putting COBOL batch on a second warehouse would have been a second security model. The constraint was: same Unity Catalog the fraud team uses, same AWS account, no “COBOL in a container on Linux” story.
Business Challenge
- Copybooks are the schema. 890 copybooks, 214 with REDEFINES, 91 with OCCURS DEPENDING ON. A naive PIC→string mapping would have corrupted packed decimal balances. COMP-3 had to become
DecimalTypewith the scale the copybook actually declared. - JCL is the DAG. 1,420 members, including PROCs with symbolic overrides. Control-M pointed at JOB names, not programs. Converting COBOL without the JCL graph would have shipped 3,800 jobs with the wrong predecessors.
- VSAM and DB2 in one night. Posting reads VSAM, enrichment hits DB2. The modern path lands both as Delta, but cutover had to keep a day where VSAM was still truth.
- 88-level conditions and REDEFINES branches. Card-type flags are 88s sitting on a REDEFINES. Emitting
ifon the wrong overlay is a silent money bug. - This was batch, not CICS. Online authorization stayed on the existing switch. Anyone selling “we modernized the mainframe” for this estate would have been lying about scope.
The MigryX Approach
Inventory hashed every program, copybook, and JCL member. Copybooks were parsed first — the schema catalog — then programs were bound to those layouts. JCL was a second graph: EXEC steps, DD datasets, COND, and PROC expansion. Control-M edges that disagreed with JCL (47 jobs) were written down, not guessed.
The COBOL emitter prefers DataFrame expressions over rdd.map. PERFORM UNTIL END-OF-FILE becomes a file read. Arithmetic on COMP-3 becomes decimal columns. Sequential updates that really are state machines (71 collections programs that walk an account history in date order) stayed as explicit window/order operations with a comment, not as “Spark will figure it out.”
JCL became Workflow YAML: one task per EXEC, parameters from symbolic overrides, failure matching COND. VSAM dumps landed in bronze Delta with the copybook schema. DB2 extracts used the same catalog names the JCL DD comments already used so ops could find “yesterday’s POST.MAST” without a decoder.
Target Architecture
z/OS batch → MigryX → Databricks on AWS (fraud lakehouse tenant)
Authorization stayed on the switch. This diagram is overnight batch only. A Workflow that claims to “replace CICS” is not this program.
Estate Inventory and Cutover Waves
| Domain | Programs | LOC | Hard parts | Wave |
|---|---|---|---|---|
| Card posting / presentment | 980 | 1.6M | COMP-3, VSAM keys | 1–2 |
| Collections / recovery | 620 | 1.1M | Ordered account walks | 2–3 |
| GL / settlement | 740 | 1.3M | Dual-run 2 closes | 3 |
| BSA / investor tapes | 510 | 0.9M | Output schema lock | 4 |
| Shared utilities | 950 | 1.3M | Called from all waves | All |
6.2M LOC is COBOL + copybooks counted once. JCL is not counted as “code.” 1,420 JCL members are the scheduler surface, not a second 6 million lines.
What we will not claim
- 81% of programs needed no human rewrite. The miss set was REDEFINES overlays, OCCURS DEPENDING ON, and the 71 ordered walks — not “COBOL is hard” as a slogan.
- Posting DAG: 6h 20m on the LPAR → 1h 35m on job clusters (same business date, 21-day dual-run). That is parallel I/O + Delta, not a 20× slide.
- $11.4M / 3 years is MIPS + OEM schedulers + the batch LPAR, minus Databricks incremental. It does not include CICS or the card switch.
Results
"We did not want COBOL running in a Linux box with a fake JES. We wanted the posting file to be a Delta table our fraud team already knew how to grant. The copybook parser was the whole game. Everything else is a Workflow."
— Head of Batch Engineering, US card issuer
Mainframe batch, Databricks runtime
Copybooks, JCL, VSAM, DB2 — parsed, not rehosted. PySpark DataFrames and Workflows, same Unity Catalog you already run.
Explore COBOL to Databricks →