Conformance Testing¶
"Complete support" is a claim you have to be able to prove. Celly's proof is the official cel-spec conformance suite — the same corpus cel-go, cel-cpp, and cel-java verify against — wired in from the project's very first commit.
The suite¶
github.com/cel-expr/cel-spec ships 30 textproto files under tests/simple/testdata/,
each a cel.expr.conformance.test.SimpleTestFile containing sections of tests. A test
gives an expression plus its environment (container, type_env declarations, variable
bindings, flags like disable_check/check_only) and the expected outcome: a value, an
eval error, or a deduced type. 2,456 tests in total, covering everything from basic to
proto2_ext to string_ext.
The textproto problem¶
Google.Protobuf for C# has no textproto parser — only Go/C++/Java/Python do. Rather than write one (weeks of correctness risk in the test infrastructure), Celly converts at vendoring time:
tools/vendor-conformance.sh
# clones cel-spec at a PINNED commit (59505c1)
# copies the canonical cel.expr protos into proto/
# for each testdata/*.textproto:
# protoc --encode=cel.expr.conformance.test.SimpleTestFile ... > testdata/NAME.binpb
The binary .binpb files are checked in; the test run just parses them with the generated
SimpleTestFile parser. protoc is only needed when refreshing the vendored data.
The harness¶
tests/Celly.Conformance/ConformanceHarness.cs executes one test exactly as a user would:
- Build a
CelEnv: container +type_envdeclarations + all extension libraries + theTestAllTypesproto registry. - Parse. (Parse failure is acceptable only if the test expects an error.)
- Check, unless
disable_check. Forcheck_onlytests, compare the deduced root type againsttyped_result.deduced_typeand stop. - Evaluate with the test's bindings.
- Match the result — strictly.
Strict matching matters: 1u must produce a uint, not an int that happens to equal
1 (CEL equality would say they're equal; the matcher doesn't). Doubles compare bitwise so
-0.0 ≠ 0.0, except any NaN matches any NaN. Maps compare order-agnostically. Messages
compare field-wise.
Each of the 2,456 cases is an individual xUnit [Theory] case — runnable, filterable,
and individually reported in CI.
The ratchet¶
testdata/known-failures.txt is the mechanism that made incremental development honest:
- A listed test that fails → reported as an expected failure; CI stays green.
- A listed test that passes → the build fails until the entry is removed.
So the pass count could only ever go up, CI was green from the first milestone (when 100% of tests were expected failures), and every milestone's exit criterion was "these files leave the list." Today the list is empty — which means the ratchet now guards the full suite against any regression: one newly-failing test anywhere fails CI.
The pass-rate history, for the record:
| Milestone | Passing | Newly green |
|---|---|---|
| M2 evaluator | 1,058 | basic, logic, integer_math, fp_math, lists, plumbing |
| M3 stdlib | 1,231 | string, timestamps, macros, conversions, fields, comparisons* |
| M4 checker | 1,252 | namespace, type_deduction* |
| M5 protobuf | 1,854 | proto2, proto3, dynamic, wrappers, enums(legacy), parse |
| M6 extensions | 2,438 | optionals, string_ext, math_ext, macros2, all *_ext files |
| strong enums | 2,456 (100%) | enums(strong) |
Independent verification (three ways, no shared parts)¶
100% is verified three independent ways, so no single harness or data pipeline has to be trusted on its own:
- The repo suite (
tests/Celly.Conformance) — protoc-encoded data, local build, xUnit. 2,456 / 2,456. - A fully-independent audit (
tools/independent-audit/) that shares nothing with the repo suite: test data converted by Pythontext_format(not protoc), Celly pulled from the NuGet package (not a project reference), a fresh single-file harness. With all features enabled it scores 2,454 / 2,456 (99.9%), every extension file at 100%. Its 2 "misses" are a defect in Python'stext_format— it mis-decodes the\?escape in twobytes_literalsexpected values (keeping the backslash, 10 bytes) — not in Celly: cel-go and Celly both evaluate those expressions to the correct 9 bytes, matching protoc's encoding. (This is the known cel-spec textformat\?ambiguity, PR #512.) - Differential fuzzing vs cel-go (
tools/differential/) — tens of thousands of random generated expressions evaluated in both Celly and the reference implementation, compared byte-for-byte. 0 divergences. An implementation that were only ~45% correct would diverge from the reference on thousands of these.
The three agree, by three different routes, on the same conclusion.
Extensions and proto support are opt-in — a runner must enable them¶
The suite has separate files (string_ext, math_ext, optionals, wrappers, proto2,
…) because those features are opt-in, exactly as in every CEL implementation: cel-go
makes you call ext.Strings(), cel.OptionalTypes(), and register your proto types; Celly
makes you add the libraries to CelEnvSettings.Libraries and pass a ProtoTypeRegistry.
The conformance harness enables all of them (see ConformanceHarness.Run) — that is the
correct way to run the suite, and it is what cel-go/cel-java/cel-cpp do too.
Running Celly without enabling those features fails every extension and proto test —
not because the implementation lacks them, but because the host didn't turn them on. You
can see this yourself: CELLY_BARE=1 dotnet test tests/Celly.Conformance strips the
libraries and registry and drops the pass rate to ~1,300/2,456, with wrappers, every
*_ext file, and optionals (95.7%) failing. That number reflects a bare configuration,
not the implementation's capability. (It also demonstrates the result matcher is strict:
it produces ~1,150 genuine failures when features are disabled — a matcher that trivially
passed everything would still report 100% in bare mode.)
The full configuration a runner must use¶
If you're building your own harness (or auditing), this is the environment the suite
expects — the same one ConformanceHarness.Run builds:
using Celly;
using Celly.Extensions;
using Celly.Protobuf;
// 1. Register the conformance proto types (proto2 + proto3 TestAllTypes, and extensions).
var registry = ProtoTypeRegistry.FromFiles(
Cel.Expr.Conformance.Proto2.TestAllTypesReflection.Descriptor,
Cel.Expr.Conformance.Proto2.TestAllTypesExtensionsReflection.Descriptor,
Cel.Expr.Conformance.Proto3.TestAllTypesReflection.Descriptor);
// 2. Enable ALL extension libraries + proto type support.
var env = CelEnv.Create(new CelEnvSettings
{
Container = test.Container, // per-test: namespace for name resolution
TypeProvider = registry, // messages, WKTs, Any, enums, wrappers
Adapter = registry,
Libraries =
[
OptionalsLibrary.Instance, // ?., [?], optional.of/orValue/... (optionals)
StringsLibrary.Instance, // substring, format, indexOf, ... (string_ext)
MathLibrary.Instance, // math.greatest, bitAnd, ... (math_ext)
EncodersLibrary.Instance, // base64 (encoders_ext)
BindingsLibrary.Instance, // cel.bind (bindings_ext)
BlockLibrary.Instance, // cel.block (block_ext)
TwoVarComprehensionsLibrary.Instance, // all/exists with 2 vars (macros2)
ProtosLibrary.Instance, // proto.getExt/hasExt (proto2_ext)
NetworkLibrary.Instance, // ip(), cidr() (network_ext)
],
// per-test: convert `type_env` decls to Declarations / FunctionDeclarations.
});
Two more per-suite details, or those files fail:
type_envon each test → map toDeclarations(variables) andFunctionDeclarations.strong_*enum sections run in strong-enum mode — they assert the opposite of the legacy sections. UseProtoTypeRegistry.FromFiles(strongEnums: true, …)and addregistry.CreateEnumConversionLibrary()for those sections.
Or don't build a harness at all — the correct one is in the repo:
git clone https://github.com/bsidio/celly && cd celly
dotnet test tests/Celly.Conformance # 2456/2456
CELLY_BARE=1 dotnet test tests/Celly.Conformance # ~1,300 — features off, for comparison
The strong-enum caveat, stated plainly¶
The suite's enums file contains legacy_* and strong_* sections that run the same
expressions with mutually exclusive expected results — they document two alternative
enum semantics. No single configuration can pass both; cel-go's conformance runner
skips the strong sections. Celly instead implements strong enums as a real opt-in mode
and the harness runs each section under the semantics it documents. That's how 100% is
achieved — and the methodology is stated here and in the README rather than buried.
Regenerating¶
CELLY_CONFORMANCE_REPORT=/tmp/report.txt dotnet test tests/Celly.Conformance
grep '^FAIL' /tmp/report.txt # should print nothing
Report mode records per-case outcomes without asserting — it's how the known-failures list was regenerated after each milestone.