Horizontal validation — results
What happens when this compiler is confronted with COBOL, specifications and
defects that were not written for it. The method is in
horizontal-validation.md; this page is the
measurement, and every number on it is read out of evidence/horizontal/.
The compiler these numbers came from
| BankLang | 0.10.0 |
| Node | v24.18.0 |
| COBOL compiler | cobc (GnuCOBOL) 3.2.0 |
| platform | darwin/arm64 |
| corpus lock | 90aaebe2d5cb |
The exact commit each lane ran on is in
evidence/horizontal/<corpus>/environment.json.
Target: IBM Enterprise COBOL 6.4. Runtime validation: GnuCOBOL. Native IBM Enterprise COBOL validation: NOT YET PERFORMED.
Nothing on this page was produced by an IBM compiler, and no result here establishes behaviour under one.
Corpora
| Corpus | Version | Licence | Redistribution |
|---|---|---|---|
| CobolCodeBench | 9d02534b7d1a |
Apache-2.0 | redistributable |
| COBOLEval | 0bb96c3114bb |
MIT | redistributable |
| X-COBOL v2 | DOI 10.5281/zenodo.14269462 | CC-BY-4.0 | derived-only |
| OpenCBS COBOL defects suite | a7a10bb0330c |
MIT | redistributable |
| NIST COBOL-85 validation suite (local) | operator-supplied | NOASSERTION | none |
Independent semantic benchmarks
CobolCodeBench
Whether BankTS can independently express a program specified by somebody else in prose, and whether the COBOL this compiler emits for it produces the byte-exact output files the benchmark expects.
| tasks discovered | 46 |
| imported into the harness | 46 |
| applicable — BankTS can express it | 19 |
| unsupported by design | 7 |
| unsupported, not yet implemented | 1 |
| the benchmark's own expectation is not derivable | 19 |
| implementations written | 20 |
| executed | 20 |
| both engines ran it | 20 |
| they agreed | 20 |
| they diverged | 0 |
ran under cobc only — never a differential pass |
0 |
| authored, of applicable | 19 / 19 (100.0%) |
| passed, of authored | 19 / 20 (95.0%) |
| passed, of applicable | 19 / 19 (100.0%) |
| passed, of all discovered | 19 / 46 (41.3%) |
Why the rest are not applicable:
| Reason | Tasks | |
|---|---|---|
| benchmark-ambiguous | 19 | task_func_01, task_func_03, task_func_08, task_func_09, task_func_10, task_func_11, task_func_12, task_func_15, task_func_19, task_func_24, task_func_26, task_func_27, task_func_29, task_func_30, task_func_31, task_func_36, task_func_39, task_func_44, task_func_55 |
| randomness | 7 | task_func_07, task_func_21, task_func_23, task_func_40, task_func_45, task_func_48, task_func_49 |
| language-gap | 1 | task_func_47 |
| Failure | Tasks |
|---|---|
| semantic-mismatch | 1 |
COBOLEval
Whether BankTS can express general algorithmic tasks against a fixed calling interface defined by somebody else, and whether the results match the benchmark's own COBOL test drivers.
| tasks discovered | 146 |
| imported into the harness | 146 |
| applicable — BankTS can express it | 0 |
| unsupported by design | 146 |
| unsupported, not yet implemented | 0 |
| the benchmark's own expectation is not derivable | 0 |
| implementations written | 0 |
| executed | 0 |
| both engines ran it | 0 |
| they agreed | 0 |
| they diverged | 0 |
ran under cobc only — never a differential pass |
0 |
| authored, of applicable | 0 / 0 |
| passed, of authored | 0 / 0 |
| passed, of applicable | 0 / 0 |
| passed, of all discovered | 0 / 146 (0.0%) |
Why the rest are not applicable:
| Reason | Tasks | |
|---|---|---|
| a fixed calling interface with no room for the transaction contract | 134 | HumanEval-1, HumanEval-100, HumanEval-101, HumanEval-102, HumanEval-103, HumanEval-104, HumanEval-105, HumanEval-106, HumanEval-107, HumanEval-108, HumanEval-109, HumanEval-11, HumanEval-110, HumanEval-113, HumanEval-114, HumanEval-116, HumanEval-117, HumanEval-118, HumanEval-119, HumanEval-12, HumanEval-120, HumanEval-121, HumanEval-122, HumanEval-123, HumanEval-124, HumanEval-126, HumanEval-127, HumanEval-128, HumanEval-13, HumanEval-130, HumanEval-131, HumanEval-132, HumanEval-134, HumanEval-135, HumanEval-138, HumanEval-139, HumanEval-14, HumanEval-140, HumanEval-141, HumanEval-142, HumanEval-143, HumanEval-144, HumanEval-145, HumanEval-146, HumanEval-147, HumanEval-148, HumanEval-149, HumanEval-15, HumanEval-150, HumanEval-152, HumanEval-153, HumanEval-154, HumanEval-155, HumanEval-156, HumanEval-157, HumanEval-158, HumanEval-159, HumanEval-16, HumanEval-160, HumanEval-161, HumanEval-162, HumanEval-163, HumanEval-17, HumanEval-18, HumanEval-19, HumanEval-23, HumanEval-24, HumanEval-25, HumanEval-26, HumanEval-27, HumanEval-28, HumanEval-29, HumanEval-3, HumanEval-30, HumanEval-31, HumanEval-34, HumanEval-35, HumanEval-36, HumanEval-39, HumanEval-40, HumanEval-41, HumanEval-42, HumanEval-43, HumanEval-44, HumanEval-46, HumanEval-48, HumanEval-49, HumanEval-5, HumanEval-51, HumanEval-52, HumanEval-53, HumanEval-54, HumanEval-55, HumanEval-56, HumanEval-57, HumanEval-58, HumanEval-59, HumanEval-6, HumanEval-60, HumanEval-62, HumanEval-63, HumanEval-64, HumanEval-65, HumanEval-66, HumanEval-67, HumanEval-68, HumanEval-69, HumanEval-7, HumanEval-70, HumanEval-72, HumanEval-73, HumanEval-74, HumanEval-75, HumanEval-76, HumanEval-77, HumanEval-78, HumanEval-79, HumanEval-8, HumanEval-80, HumanEval-82, HumanEval-83, HumanEval-84, HumanEval-85, HumanEval-86, HumanEval-88, HumanEval-89, HumanEval-9, HumanEval-90, HumanEval-91, HumanEval-93, HumanEval-94, HumanEval-96, HumanEval-97, HumanEval-98 |
| COMP-2 / COMP-1 | 12 | HumanEval-0, HumanEval-133, HumanEval-151, HumanEval-2, HumanEval-20, HumanEval-21, HumanEval-45, HumanEval-47, HumanEval-71, HumanEval-81, HumanEval-92, HumanEval-99 |
Real-world coverage
X-COBOL v2
No behavioural oracle. These are files, not tests: nothing here can establish that anything computes the right answer, and a representability figure is a statement about language scope rather than about correctness.
| COBOL files discovered | 5195 |
| read without error | 5195 / 5195 (100.0%) |
| analyser failures | 0 |
| Representability | Files |
|---|---|
| fully-representable | 1543 / 5195 (29.7%) |
| representable-with-adaptation | 2730 / 5195 (52.6%) |
| unsupported-by-design | 533 / 5195 (10.3%) |
| unsupported-not-yet-implemented | 331 / 5195 (6.4%) |
| analyser-failure | 0 / 5195 (0.0%) |
| unknown | 58 / 5195 (1.1%) |
Constructs BankTS cannot express, ranked by how often they occur:
| Construct | Support | Files | Share |
|---|---|---|---|
go-to |
adaptation | 2473 | 47.6% |
perform-thru |
adaptation | 2394 | 46.1% |
reference-modification |
adaptation | 600 | 11.5% |
usage-index |
adaptation | 452 | 8.7% |
string-unstring |
adaptation | 393 | 7.6% |
inspect |
adaptation | 269 | 5.2% |
usage-pointer |
unsupported-by-design | 244 | 4.7% |
file-relative |
unsupported-not-yet-implemented | 192 | 3.7% |
screen-section |
unsupported-by-design | 138 | 2.7% |
external-data |
unsupported-not-yet-implemented | 110 | 2.1% |
copy-replacing |
unsupported-not-yet-implemented | 87 | 1.7% |
entry-point |
unsupported-by-design | 79 | 1.5% |
alter |
unsupported-by-design | 70 | 1.3% |
comp-float |
unsupported-by-design | 37 | 0.7% |
go-to-depending |
unsupported-by-design | 28 | 0.5% |
The rules behind each verdict are in packages/horizontal-validation/src/representability.ts, one row per construct. countOf lowers to INSPECT TALLYING FOR ALL and replaceChars to INSPECT CONVERTING. REPLACING, and the BEFORE/AFTER ranges, have no BankTS form — 780 and 625 statements in the corpus respectively.
OpenCBS COBOL defects suite
Defects reconstructed from public forum posts, so they over-represent what people ask about rather than what most often reaches production. A defect BankTS cannot express at all is prevented trivially and is recorded as such rather than counted as a save.
| COBOL files discovered | 53 |
| read without error | 53 / 53 (100.0%) |
| analyser failures | 0 |
| Representability | Files |
|---|---|
| fully-representable | 22 / 53 (41.5%) |
| representable-with-adaptation | 30 / 53 (56.6%) |
| unsupported-by-design | 0 / 53 (0.0%) |
| unsupported-not-yet-implemented | 0 / 53 (0.0%) |
| analyser-failure | 0 / 53 (0.0%) |
| unknown | 1 / 53 (1.9%) |
Constructs BankTS cannot express, ranked by how often they occur:
| Construct | Support | Files | Share |
|---|---|---|---|
go-to |
adaptation | 25 | 47.2% |
usage-index |
adaptation | 4 | 7.5% |
string-unstring |
adaptation | 3 | 5.7% |
reference-modification |
adaptation | 1 | 1.9% |
perform-thru |
adaptation | 1 | 1.9% |
The rules behind each verdict are in packages/horizontal-validation/src/representability.ts, one row per construct. countOf lowers to INSPECT TALLYING FOR ALL and replaceChars to INSPECT CONVERTING. REPLACING, and the BEFORE/AFTER ranges, have no BankTS form — 780 and 625 statements in the corpus respectively.
What each language change moved
X-COBOL representability before and after each feature, over the same
5,195 files. The before column is the measurement as it stood on the
commit named in evidence/horizontal-history/index.json.
inspect-classification
Measured against 216de338cc15.
| Verdict | Before | After | Change |
|---|---|---|---|
| fully-representable | 1543 | 1543 | 0 |
| representable-with-adaptation | 2548 | 2730 | +182 |
| unsupported-by-design | 533 | 533 | 0 |
| unsupported-not-yet-implemented | 513 | 331 | -182 |
| analyser-failure | 0 | 0 | 0 |
| unknown | 58 | 58 | 0 |
line-sequential
Measured against 8abf3da6ebc5.
| Verdict | Before | After | Change |
|---|---|---|---|
| fully-representable | 1388 | 1543 | +155 |
| representable-with-adaptation | 2478 | 2730 | +252 |
| unsupported-by-design | 533 | 533 | 0 |
| unsupported-not-yet-implemented | 738 | 331 | -407 |
| analyser-failure | 0 | 0 | 0 |
| unknown | 58 | 58 | 0 |
Defect benchmark
41 reconstructed defects. 9 are prevented at compile time by a BankTS program the compiler refuses; see the matrix.
COBOL conformance
NIST COBOL-85 is never downloaded or redistributed by this repository. No local copy was supplied, so this lane is unavailable — which is not the same as passing.