Most mainframe estates that come to us have the same shape. There are thousands of COBOL batch programs, hundreds of copybooks, JCL job streams that have run nightly for decades, and DB2 tables at either end. The business logic is sound. What the owners want to escape is the MIPS bill, the shrinking pool of people who can change it safely, and a batch window that no longer fits.
Databricks is a natural landing zone for that batch work. Record-at-a-time COBOL becomes set-based PySpark. Flat files become Delta tables governed in Unity Catalog. JCL becomes Lakeflow Jobs (formerly Databricks Workflows). Getting there safely depends on three layers that generic code translators tend to skip: the bytes, the arithmetic and the control flow.
What moves, and what doesn't
Be precise about scope from day one. Batch has clear inputs, outputs and schedules, and it maps cleanly onto a lakehouse. Online transaction processing is a different problem with a different answer.
In scope: mainframe batch
- COBOL batch programs
- Copybooks and record layouts
- JCL job streams, PROCs, SORT steps
- Embedded DB2 SQL
- Sequential and GDG datasets
Not in this program
- CICS and other online transactions
- Green-screen (3270) applications
- Rehosting COBOL on an emulator
This is a modernization, not a rehost: the output is native PySpark and Delta, not COBOL running on rented hardware. Your team keeps and maintains that code.
A batch job, before and after
Here is a typical monthly job: sort the account master, run the interest program, load the result into DB2. On the mainframe that takes three JCL steps, two datasets and a generation data group (GDG). On Databricks it becomes two tasks, a condition check and three Delta tables.
The bytes: copybooks, EBCDIC and packed decimal
A mainframe file has no header, no delimiters and no types. The copybook is its only schema, and it describes bytes, not columns. Before any business logic runs, every record has to be decoded exactly as the COBOL program would read it.
Copybook ACCTREC.cpy 01 ACCT-REC.
05 ACCT-ID PIC X(10).
05 ACCT-BAL PIC S9(9)V99 COMP-3.
05 ACCT-TYPE PIC X(01).
88 SAVINGS VALUE 'S'.
88 CHECKING VALUE 'C'.
05 OPEN-DATE PIC 9(08).
05 INT-RATE PIC S9(2)V9(4) COMP-3.
Four details break most homegrown readers:
- EBCDIC is not ASCII. The letter A is
x'C1', notx'41', and the code page (037, 1140, 273 and so on) changes characters such as$,£and accented letters. - Packed decimal (COMP-3) stores two digits per byte plus a sign nibble. Converting the file to ASCII before decoding destroys those bytes.
- Zoned decimal signs are overpunched into the last digit:
x'C5'means +5,x'D5'means −5. In a text viewer they appear as letters. - REDEFINES and OCCURS DEPENDING ON mean the layout depends on the data: a record-type byte picks the layout, and a count field sets the length.
On Databricks, the decoded result lands in a Bronze Delta table with real types. One common route is an open-source Spark reader for COBOL data such as Cobrix, which applies the copybook directly:
PySpark · bronze ingestraw = (spark.read.format("cobol")
.option("copybook", "/Volumes/mainframe/raw/copybooks/ACCTREC.cpy")
.option("record_format", "F")
.option("ebcdic_code_page", "cp037")
.load("/Volumes/mainframe/raw/landing/ACCT.MASTER.D260928"))
(raw.write.format("delta").mode("overwrite")
.saveAsTable("mainframe.bronze.acct_master"))
Whichever reader you use, the test is the same: decoded values must match what the COBOL program saw, field by field, including sign and scale, before any logic is converted.
The arithmetic: COBOL programs to PySpark
A COBOL batch program reads a record, applies rules and writes a record, inside a PERFORM UNTIL loop. Rewritten as a DataFrame, the same rules run across the whole file in parallel, and that is where batch windows shrink. The rules themselves need care, because COBOL arithmetic has its own conventions.
01 WS-INTEREST PIC S9(7)V99 COMP-3.
...
PERFORM UNTIL END-OF-FILE
READ ACCT-FILE INTO ACCT-REC
AT END SET END-OF-FILE TO TRUE
END-READ
IF NOT END-OF-FILE AND SAVINGS
COMPUTE WS-INTEREST = ACCT-BAL * INT-RATE / 1200
WRITE INT-REC FROM WS-INT-REC
END-IF
END-PERFORM.
PySpark · step020_acctint
from pyspark.sql import functions as F
accts = spark.table("mainframe.bronze.acct_master")
raw = F.col("acct_bal") * F.col("int_rate") / F.lit(1200)
# COBOL COMPUTE without ROUNDED truncates toward zero to the receiving scale
truncated = F.when(raw >= 0, F.floor(raw * 100)).otherwise(F.ceil(raw * 100)) / 100
interest = (accts
.filter(F.col("acct_type") == "S") # 88 SAVINGS
.withColumn("ws_interest", truncated.cast("decimal(9,2)"))
.select("acct_id", "acct_bal", "int_rate", "ws_interest"))
interest.write.format("delta").mode("overwrite").saveAsTable("finance.silver.acct_interest")
COMPUTE without ROUNDED truncates. Spark's decimal cast rounds. On a million accounts, that one-cent difference per row is a reconciliation failure the day after cutover. Overflow behaves differently too. If the result exceeds S9(7)V99, COBOL silently drops the high-order digits unless the program coded ON SIZE ERROR. Spark returns NULL or raises an error, depending on ANSI mode. Neither matches COBOL by default, so each case should be flagged and decided on purpose, not inherited by accident. Also note that the JCL SORT step is gone: the filter covers its INCLUDE, and the order no longer matters because nothing downstream depends on it.Other COBOL patterns follow the same principle of keeping the result and changing the mechanics:
- Control-break processing, which compares each record's key with the previous one to write subtotals, becomes a
groupByor a window function. - EVALUATE / IF chains and 88-level conditions become
F.when()chains with named conditions. - Table lookups (
SEARCH ALLover an OCCURS table loaded at start-up) become broadcast joins. - DB2 cursors (
DECLARE/FETCHloops) become one Spark SQL query against Delta.
The control flow: JCL to Lakeflow Jobs
JCL defines order, datasets and conditions. Its COND parameter is famously back to front: COND=(4,LT) means skip this step if 4 is less than any earlier return code. In other words, it runs only when every prior step ended with RC ≤ 4. Get the logic backwards and a failed step quietly lets the load run.
//ACCTINT JOB (ACCT),'MONTHLY INTEREST',CLASS=A
//STEP010 EXEC PGM=SORT
//SORTIN DD DSN=PROD.ACCT.MASTER,DISP=SHR
//SORTOUT DD DSN=&&SORTED,DISP=(NEW,PASS)
//SYSIN DD *
SORT FIELDS=(1,10,CH,A)
INCLUDE COND=(17,1,CH,EQ,C'S')
/*
//STEP020 EXEC PGM=ACCTINT,COND=(4,LT)
//ACCTIN DD DSN=&&SORTED,DISP=(OLD,DELETE)
//INTOUT DD DSN=PROD.ACCT.INTEREST(+1),DISP=(NEW,CATLG)
//STEP030 EXEC PGM=IKJEFT01,COND=(0,NE)
//SYSTSIN DD *
DSN SYSTEM(DB2P)
RUN PROGRAM(ACCTLOAD) PLAN(ACCTPLN)
/*
Databricks Asset Bundle · resources/acct_monthly_interest.yml
resources:
jobs:
acct_monthly_interest:
name: acct_monthly_interest
tasks:
- task_key: step020_acctint
notebook_task:
notebook_path: ../src/acctint.py
- task_key: rc_check
depends_on:
- task_key: step020_acctint
condition_task:
op: EQUAL_TO
left: "{{tasks.step020_acctint.values.return_code}}"
right: "0"
- task_key: step030_acctload
depends_on:
- task_key: rc_check
outcome: "true"
notebook_task:
notebook_path: ../src/acctload.py
The converted program publishes its return code with dbutils.jobs.taskValues.set("return_code", rc), so warning-level outcomes (RC 4) are handled explicitly rather than lost. Other JCL constructs map just as directly:
- GDG generations.
(+1),(0)and(-1)become Delta table versions: write a new version, read the current one, or read the previous one withVERSION AS OF. - Temporary datasets (
&&) usually disappear into DataFrames and are never written. - DFSORT steps become
orderBy,filterordropDuplicates. ForSUM FIELDS=NONE, which record survives among duplicates is only guaranteed withEQUALS, so check which one the business relies on. - Restart and checkpoint logic becomes Lakeflow Jobs retries and repair runs.
- Scheduler calendars (Control-M, CA-7, TWS) become job schedules and triggers.
Artifact mapping
| Mainframe artifact | On Databricks | Watch for |
|---|---|---|
| Copybook | Spark schema, Bronze Delta table | Code page, COMP-3, zoned signs, REDEFINES, ODO |
| COBOL batch program | PySpark notebook or job task | Truncation vs rounding, SIZE ERROR, intermediate precision |
| 88-level conditions | F.when() predicates | Multiple values and THRU ranges |
| Control breaks | groupBy / window functions | Key ordering, last-group flush |
| Embedded DB2 SQL | Spark SQL on Delta | Cursor loops, NULL indicators, DB2 date types |
| JCL steps and COND | Lakeflow Jobs tasks, condition tasks | Inverted COND logic, RC 4 warnings |
| GDG datasets | Delta versions / time travel | Retention limits vs GDG LIMIT |
| DFSORT / ICETOOL | orderBy, filter, dropDuplicates | Duplicate survivorship without EQUALS |
| Scheduler (Control-M, CA-7) | Job schedules and triggers | Calendars, cross-job dependencies |
How a mainframe migration runs
Start with two or three job streams that share datasets. That is enough to exercise copybook decoding, program logic, DB2 access and JCL control flow together, and small enough to finish in a quarter. During the parallel run, both platforms process the same inputs and the outputs are compared record by record and field by field: sign, scale, dates and all. Cutover happens when that comparison is clean, not when the code compiles.
The proof is the product. Operations and audit teams do not sign off on converted code. They sign off on evidence that the output matches. Plan the parity report as a deliverable from the first week.
Key takeaways
- Scope mainframe batch (COBOL, copybooks, JCL, DB2 SQL) separately from online and CICS work.
- Decode bytes first. EBCDIC code pages, COMP-3 and zoned signs must match what the COBOL program saw before any logic moves.
- COBOL truncates where Spark rounds, and overflows silently where Spark returns NULL or fails. Decide each case on purpose.
- JCL
CONDlogic is inverted. It maps to explicit Lakeflow Jobs condition tasks driven by return codes. - GDGs become Delta versions, temporary datasets disappear, and many SORT steps collapse into one DataFrame operation.
See one of your job streams on Databricks
Bring a JCL job with its COBOL programs and copybooks. We'll walk through the converted PySpark and Lakeflow Job, and how the output is checked against your mainframe run.
Book a demo COBOL to Databricks