Elwood is a standalone JSON transformation DSL inspired by KQL pipes, LINQ lambdas, and JSONPath navigation. It is designed for high performance, rich error reporting, and cross-platform use.
The .NET implementation is the reference engine. A TypeScript implementation provides the same language in browsers, Node.js, and edge runtimes. Both implementations share a conformance test suite to guarantee identical behavior.
Goal: A working JSON transformation engine.
- Solution structure: Elwood.Core, Elwood.Json, Elwood.Newtonsoft (placeholder), Elwood.Cli
- Lexer and recursive descent parser
- AST with 30+ node types
- Tree-walking evaluator
- IElwoodValue abstraction (decoupled from JSON library)
- System.Text.Json adapter (JsonNode)
- JSONPath navigation:
$,$.field,$[*],$[0],$[2:5],$[-2:],$..field - Auto-mapping: property access maps over arrays
- KQL-style pipe operators:
where,select,selectMany,orderBy,groupBy,distinct,take,skip,batch,join(inner/left/right/full),concat,index,reduce,count,sum,min,max,first,last,any,all,match - Named lambda expressions:
x => expr,(acc, x) => expr - Implicit
$context in pipe operations -
letbindings (script-scoped variables) -
memofunctions (memoized lambdas for expensive computations) -
if/then/elseconditionals - Pattern matching:
| match "val" => result, _ => default - Spread operator:
{ ...obj, newProp: val } - String interpolation:
`Hello {$.name}` - Arithmetic, boolean, and comparison operators (including string comparison)
- 70+ built-in methods (string, numeric, datetime, crypto, null-checks, etc.)
- Rich error reporting with line/column, "Did you mean?" suggestions
- CLI tool: REPL, eval, run modes, stdin pipe support
- 90 tests (65+ file-based with explanations, 25 code-based)
- Documentation: syntax reference, changelog, 65 explanation files as tutorials
Goal: Make Elwood available in browsers, Node.js, Deno, Bun, and edge runtimes (Cloudflare Workers, Vercel Edge) via a TypeScript implementation that is behaviorally identical to the .NET reference engine.
Detailed implementation plan: See docs/typescript-port-plan.md
-
LazyArrayValuewrapsIEnumerable<IElwoodValue>without materializing; streaming operators (where,select,selectMany,take,skip,distinct,index) return lazy arrays -
$[*]wildcard returns lazy array (critical for short-circuit) - Only materializing operators call
.ToList():orderBy,groupBy,count,sum,min,max,last,batch,join,reduce -
take(n)short-circuits —take(1)on 100K items: 19ms vs 400ms+ before -
firstshort-circuits without materializing - Pipeline endpoints materialize via
ToConcreteValue()for final JSON output - 4 benchmark tests with results logged to
tests/Elwood.Core.Tests/Benchmarks/results.log(last 10 runs) - All 95 tests pass (91 conformance + 4 benchmarks)
- All 48 explanation files cleaned — legacy system references replaced with generic terms
- Sample data and examples updated
- Source code comments cleaned
- 68 test cases extracted to
spec/test-cases/{name}/(per-directory: script.elwood, input.json, expected.json, explanation.md) - .NET code moved under
dotnet/(dotnet/src/,dotnet/tests/,dotnet/Elwood.slnx) - FileBasedTests.cs updated to discover from
spec/test-cases/ - CLAUDE.md updated with new structure
- 97/97 tests pass after restructure
- Initialize TS project (
ts/) with TypeScript, Vitest, zero runtime dependencies - Conformance test runner that reads from
../../spec/test-cases/— 66 tests discovered, all fail with "Not implemented" - Port lexer (28 unit tests passing)
- Port parser (24 unit tests passing)
- Port evaluator (tree-walk, scope management, auto-mapping, JSONPath navigation)
- Port built-in functions — 70+ methods:
- Array/pipe operators (where, select, orderBy, groupBy, join, reduce, etc.)
- String methods (toLower, replace, split, urlEncode, sanitize, etc.)
- Numeric methods (round, floor, ceiling, convertTo, etc.)
- DateTime methods (now, dateFormat, dateAdd, toUnixTimeSeconds, etc.)
- Crypto methods (hash via node:crypto MD5, rsaSign via node:crypto RSA-SHA1+PKCS1)
- Null-check, object, and type methods
- Error reporting with line/column and "Did you mean?" suggestions
- All 66 conformance tests passing (122 total: 28 lexer + 26 parser + 67 conformance + 1 discovery)
New features follow the workflow: spec test case first → implement in .NET → implement in TS → update docs. CI runs both conformance suites on every PR.
Goal: Open-source release of Elwood with both implementations, CI, and package distribution.
Depends on: Phase 1b complete (repo is in final structure, both implementations pass conformance suite).
- Published to GitHub: https://github.com/max-favilli/elwood
- README.md with quick start, examples, CLI download links, playground screenshot
- LICENSE file (MIT)
- GitHub Actions CI: build + test on push (both .NET and TS) —
.github/workflows/ci.yml - Conformance gate: both implementations must pass all
spec/test-cases/ - Automated release workflow (
.github/workflows/release.yml) — triggered on version tags - v0.1.0 released with native binaries for Windows, macOS, Linux
- NuGet package metadata (Elwood.Core, Elwood.Json — multi-target net8.0;net10.0)
-
global.jsonpinning SDK version -
dotnet toolconfigured (dotnet tool install --global Elwood.Cli→elwoodcommand) - Native AOT binaries on GitHub Release (linux-x64, macos-x64, win-x64)
- Publish NuGet packages to nuget.org —
Elwood.Core,Elwood.Json,Elwood.Pipeline,Elwood.Cli,Elwood.Xlsx,Elwood.Parquet(v0.3.0)
- npm package published:
@elwood-lang/coreon npmjs.com - Browser compatibility documented (crypto limitation in known-issues.md)
- Browser-based interactive playground: https://max-favilli.github.io/elwood/
- Monaco Editor with Elwood syntax highlighting, autocomplete, snippet placeholders
- Example gallery (68 examples from
spec/test-cases/) with search, categories, rendered markdown explanations - File loading (local file picker + drag-and-drop, stays in browser)
- Shareable links via lz-string URL compression
- GitHub Pages deployment via GitHub Actions (
.github/workflows/deploy-playground.yml) - Spec:
docs/playground-spec.md
-
Elwood.Apiproject — minimal ASP.NET API:POST /api/evaluate+/health - Dockerfile (multi-stage, Alpine-based)
- GitHub Actions workflow to build and push container on release tags (
.github/workflows/docker.yml) - Usage:
docker run -p 8080:8080 ghcr.io/max-favilli/elwood-api - Enables integration with any iPaaS via HTTP
-
iterate(seed, fn)andtakeWhile— lazy sequence generation -
.parseJson()— deserialize embedded JSON strings - Identify and port any remaining common JSON transformation functions not yet covered
Goal: Use pure Elwood scripts (.elwood files) as the transformation format, with format conversion for non-JSON inputs/outputs. No YAML maps.
An earlier design proposed YAML documents where the tree structure mirrors the output shape and leaf values are Elwood expressions. This was abandoned for two reasons:
- Deep indentation — A 20-level deep output (common in SAP IDocs, complex XML) means 40+ spaces of indent before the content. YAML's indentation-as-structure makes depth visible and painful.
- Two syntaxes in one file — YAML for structure + Elwood for values means no editor can fully syntax-highlight, autocomplete, or error-check both. Every YAML-embedded-DSL (GitHub Actions + bash, Helm + Go templates) suffers from this — the embedded language gets zero tooling.
Pure Elwood scripts solve both problems:
- One syntax → full editor support (highlighting, autocomplete, error reporting)
letdecomposition flattens deep nesting into named, testable piecesmemohandles repeated patterns (e.g., 82 SAP segments → 30 one-liner function calls)- Already works today with
elwood run script.elwood --input data.json
// artmas09.elwood — replaces a 6500-line JSON map
let ausprt = memo (src, charName) => {
"@SEGMENT": "1", FUNCTION: "009", CHAR_NAME: charName,
MATERIAL_LONG: $.SAP_STUFF.MATERIAL_LONG,
CHAR_VALUE_LONG: src.CHAR_VALUE_LONG
}
let ediDc40 = { "@SEGMENT": "1", TABNAM: "EDI_DC_40", MANDT: "400", ... }
let matHead = { "@SEGMENT": "1", MATL_TYPE: $.SAP_STUFF.MATL_TYPE, ... }
return {
ARTMAS09: {
IDOC: {
"@BEGIN": "1",
EDI_DC40: ediDc40,
E1BPE1MATHEAD: matHead,
E1BPE1AUSPRT: [
ausprt($.C8_PRODUCTSUBTYPE, "C8_PRODUCTSUBTYPE"),
ausprt($.JW_PRODUCTGROUP, "JW_PRODUCTGROUP"),
// ... 30 more one-liners instead of 82 copy-pasted map nodes
]
}
}
}
Elwood operates on JSON internally. Format converters handle non-JSON inputs/outputs:
any format ──→ [input conversion] ──→ JSON ──→ Elwood script ──→ JSON ──→ [output conversion] ──→ any format
| Format | Input (→ JSON) | Output (JSON →) |
|---|---|---|
| JSON | Native (no conversion) | Native (no conversion) |
| CSV | Rows → array of objects | Array of objects → rows |
| XML | Elements → objects, attributes → @ properties |
Objects → elements |
| XLSX | Sheet → array of objects | Array of objects → sheet |
| Text | Content as string or line-split array | Join array / render string |
Two ways to use format conversion:
CLI flags — for simple cases where the entire input is one format and the entire output is another:
elwood run transform.elwood --input data.csv --input-format csv
elwood run transform.elwood --input data.xml --output result.csv --output-format csvIn-script functions — for full control within the script (custom options, mixed formats, converting mid-pipeline):
// Parse CSV with custom delimiter
let orders = $.rawCsv | fromCsv({ delimiter: ";", headers: true })
// Transform
let result = orders | where(.amount > 100) | select({ id: .id, total: .amount })
// Output as CSV
return result | toCsv()
// Parse XML, transform, output as JSON
let items = $.xmlPayload | fromXml()
return items.catalog.products | select({ sku: .@id, name: .title })
// Mix formats in one script
let products = $.csvData | fromCsv({ headers: true })
let categories = $.xmlCategories | fromXml()
return products | select(p => {
...p,
categoryName: categories.list[*] | first(c => c.@id == p.catId) | select(.name)
})
| Function | Description | Options | Status |
|---|---|---|---|
fromCsv(options?) |
Parse CSV string → array of objects | delimiter, headers, quote, skipRows, parseJson |
✅ |
toCsv(options?) |
Array of objects → CSV string | delimiter, headers, alwaysQuote |
✅ |
fromXml(options?) |
Parse XML string → JSON object | attributePrefix, stripNamespaces |
✅ |
toXml(options?) |
JSON object → XML string | rootElement, attributePrefix, declaration |
✅ |
fromXlsx(options?) |
Parse XLSX (base64) → array of objects | headers, sheet Extension |
✅ |
toXlsx(options?) |
Array of objects → XLSX (base64) | headers, sheet Extension |
✅ |
fromParquet(options?) |
Parse Parquet (base64) → array of objects | Extension | ✅ |
toParquet(options?) |
Array of objects → Parquet (base64) | schema, compression Extension, .NET only |
✅ |
fromText(options?) |
Split text into lines or structured data | delimiter (default \n) |
✅ |
toText(options?) |
Join array into text | delimiter (default \n) |
✅ |
- Built-in format functions:
fromCsv,toCsv,fromText,toText - Built-in format functions:
fromXml,toXml - CLI
--input-formatand--output-formatflags (auto-detect from file extension, or explicit override) - CLI integration tests (13 tests: eval, run, format detection, output conversion, stdin, error handling)
- Format converters — CSV:
- Configurable delimiter, headers, quote character
-
skipRows— skip metadata/title rows before data -
headers: false— auto-generated alphabetic column names (A, B, C, ..., AA, AB) -
parseJson: true— auto-detect and deserialize JSON values in cells -
alwaysQuote— force-quote all fields in toCsv output
- Format converters — Text: line splitting/joining with configurable delimiter
- Format converters — XML (attributes →
@prefix, repeated elements → arrays, namespace stripping) - Format converters — XLSX (sheet selection, header row) — via separate extension packages (
Elwood.Xlsx,@elwood-lang/xlsx) - Extension/plugin API:
RegisterMethod(.NET),registerMethod(TS) for optional packages - Format converters — Parquet (read/write via extension:
Elwood.Parquet,@elwood-lang/parquet) - Binary pass-through: CLI
--input-format binaryreads files as base64 - Multi-format test inputs: test runners support
input.csv,input.txt,input.xml - Parser fix:
$.method()resolves correctly when$is a non-object value -
.parseJson()general-purpose method +fromCsv({ parseJson: true })convenience option - Bracket property access:
obj["@attr"]for XML attributes and special-character keys -
.first()/.last()on strings return first/last character - Extension exceptions wrapped as diagnostics (not raw crashes)
- 86 conformance test cases, 137 .NET tests (115 core + 15 CLI + 7 Parquet), 144 TS tests (140 core + 4 XLSX)
Goal: Ensure Elwood outperforms comparable JSONPath-based transformation engines.
Benchmarked in-process on 100K rows (fair comparison, same machine, no HTTP overhead):
| Test | Elwood .NET | Legacy baseline | Speedup |
|---|---|---|---|
where active | select name |
121ms | 240ms | 2.0x faster |
select with toString + charArray concat |
836ms | 1,819ms | 2.2x faster |
- Lazy evaluation via
LazyArrayValuestreams items through pipeline stages without materializing intermediate arrays - LINQ integration — .NET's optimized
Where/Select/TakeonIEnumerableavoids per-item allocation - The JIT compiler already optimizes the hot interpreter loop after warm-up
Expression Tree compilation was implemented and tested but provided no speedup over the interpreter. The interpreter's lazy streaming is more efficient than compiled fused loops that materialize arrays. The compiled mode was removed to reduce complexity.
Benchmarked on 100K rows — the TS interpreter is already ~5x faster than .NET:
| Test | .NET | TypeScript |
|---|---|---|
| where+select name | 121ms | 24ms |
| toString + charArray concat | 836ms | 173ms |
V8's JIT aggressively optimizes native array methods (filter, map). Generator-based lazy evaluation is unnecessary — the eager approach with V8-optimized array ops is faster.
- Bypass
IElwoodValueabstraction for directJsonNodeaccess in .NET hot paths — only if 100MB+ workloads need it
Goal: Use YAML to describe complete data integration pipelines — sources, transformations, and destinations — in a single document. YAML handles declarative orchestration (triggers, connections, destinations); Elwood scripts (.elwood files) handle transformation logic.
Real-world integration configs embed complex expressions in many YAML values — multi-line filter chains, conditional logic, method chains, etc. This creates the same two-syntax-in-one-file problem we rejected for transformation maps.
Solution: YAML for structure, external .elwood scripts for any non-trivial expression.
Inline in YAML (OK):
- Static values:
trigger: http,concurrency: 100,container: output - Simple paths:
$.request.season,$.code - Short interpolation:
`{$.request.season}-images`
External .elwood file (required for):
- Anything with pipes (
|) - Conditionals (
if/then/else) - Method chains (
.toUpper().split()...) - Filters (
where,in) - Multi-line logic
Guideline (not enforced): simple $.field, short interpolation, or brief expressions stay inline. Complex logic with multiple pipes, conditionals, or long method chains goes in an external .elwood file. This is a recommendation — short pipes like $.items[*] | take(5) are fine inline. A future elwood validate command could warn (not error) when inline expressions exceed a complexity threshold.
# pipeline.elwood.yaml
version: 2
sources:
- name: api-trigger
trigger: http
endpoint: /api/data/{category}
contentType: json
map: request-map.elwood # ← external script
- name: file-source
trigger: pull
from:
fileShare:
connectionString: ${FILE_SHARE_CONN}
path: /{$.request.category}/data # ← simple interpolation, OK inline
map: source-transform.elwood # ← external script
join:
path: $
keys: []
outputs:
- name: publish-to-fileshare
path: filter-results.elwood # ← complex filter → external script
outputId: output-id.elwood # ← method chains → external script
contentType: $.contentType # ← simple path, OK inline
concurrency: 100 # ← static value
map: output-map.elwood # ← external script (reusable!)
destinations:
fileShare:
- connectionString: ${FS_CONN}
filename: output-filename.elwood # ← complex interpolation → external
- name: publish-to-sftp
path: filter-results.elwood # ← same filter, reused!
outputId: output-id.elwood # ← same ID logic, reused!
contentType: $.contentType
concurrency: 50
map: output-map.elwood # ← same map, reused!
destinations:
sftp:
- connectionString: ${SFTP_CONN}
filename: output-filename.elwood # ← same filename logic, reused!External scripts are reusable — the same output-id.elwood and filter-results.elwood are shared across multiple outputs. They're also independently testable via the CLI.
| In YAML (inline) | Inline Elwood (simple) | External .elwood (complex) |
|---|---|---|
trigger: http |
path: $.results[*] |
path: filter-active.elwood |
contentType: json |
contentType: $.contentType |
outputId: generate-id.elwood |
concurrency: 100 |
endpoint: /api/{$.category} |
map: transform.elwood |
container: output |
filename: /{$.code}.json |
filename: build-path.elwood |
Guideline: static config → plain YAML. Simple expressions → inline. Complex logic → external .elwood file. This is a best practice, not enforced. Short inline pipes are fine when readable.
IDM (Intermediate Data Model): shared JSON document built progressively by sources, consumed by outputs. Sources → IDM → Outputs → Destinations.
Script bindings: named root variables set by the executor, available in .elwood scripts:
| Binding | Available in | What it contains |
|---|---|---|
$ |
Source maps | Raw source payload |
$ |
Output maps | Current fan-out slice |
$source |
Everywhere | Source metadata (trigger, headers, eventId) |
$idm |
Source maps (after first source) | Current IDM state from previous sources |
$idm |
Output maps | Complete IDM |
$output |
Output maps | Full array from path (all slices) |
$secrets |
YAML properties | Secret references loaded from provider |
Scripts that don't need metadata work identically in the playground and in a pipeline — $ is always the data.
Fan-out: both sources and outputs support path — slices the IDM, processes once per slice with optional concurrency.
Full schema reference: docs/pipeline-yaml-reference.md
The current model is strictly sources → IDM → outputs: all sources complete before any output runs. This works well for most integrations, but some real-world flows need to deliver data mid-pipeline — for example, archiving the raw payload to blob storage before enrichment, or calling an API and delivering its response to a file THEN using that result to call a second API.
Today, these cases require splitting into multiple chained pipelines (pipeline A outputs to a queue, pipeline B triggers from the queue). This works but spreads a single logical flow across multiple configs, making it harder to reason about, test, and monitor end-to-end.
A staged execution model would allow interleaving sources and outputs within a single pipeline:
stages:
- sources:
- name: trigger
trigger: http
outputs:
- name: archive-raw
destinations:
blob:
- container: raw-payloads
- sources:
- name: enrich
trigger: pull
from:
http:
url: https://api.example.com/enrich
body: $.request
outputs:
- name: api-response
response: trueEach stage runs in order. Within a stage, sources build the IDM, then outputs consume it. The IDM accumulates across stages. The key distinction between sources and outputs is preserved: a source fetches from ONE place and writes to the IDM; an output delivers from the IDM to ONE OR MORE destinations (fan-out on the delivery side).
Status: deferred. The current sources→outputs model handles the common case. If real-world integrations repeatedly require chaining multiple pipelines for what is logically one flow, this design will be revisited. Evidence from actual pipeline authoring will inform the decision.
| Executor | Purpose | Sources | Destinations |
|---|---|---|---|
| CLI Executor | Development + testing with saved payloads | Local files (one per named source) | Local files |
| Sync Executor | End-to-end local execution, connects to real sources | HTTP calls, file shares, queues | Real destinations |
| Cloud Executors (Azure, AWS) | Production, distributed, async | Triggers + pull sources | Real destinations |
All three share the same pipeline parser, script resolver, and transformation engine. They differ only in how they acquire source data and deliver outputs.
CLI Executor usage:
# Single source
elwood pipeline run pipeline.elwood.yaml --source api-trigger=payload.json
# Multi-source — provide envelope files with source metadata
elwood pipeline run pipeline.elwood.yaml \
--source api-trigger=trigger-envelope.json \
--source product-api=product-response.json
# Outputs written to local files (stdout or --output-dir)Envelope file format (for CLI executor):
{
"source": {
"name": "api-trigger",
"trigger": "http",
"eventId": "evt-abc-123",
"http": { "method": "POST", "headers": { "X-Correlation-Id": "corr-789" } }
},
"payload": {
"orders": [{ "id": 1, "active": true }]
}
}The executor splits it: $ = envelope.payload, $source = envelope.source. Plain JSON files (no envelope) are also accepted — $ = the file content, $source = minimal defaults.
Step 1 — Pipeline YAML schema + parser: ✅
- Define integration YAML schema (sources, outputs, destinations)
-
Elwood.Pipelineproject — YAML parser (using YamlDotNet) - Resolve
.elwoodfile references relative to YAML file location - Source envelope schema (source metadata + payload)
- PipelineExecutor: source maps → IDM → output path → output maps
-
depends— source dependency graph + stage resolution -
pathfan-out on sources (with$slicebinding) -
$source,$idm,$outputbindings in evaluator -
$-prefixed identifiers in both .NET and TS lexers -
$secretsresolution from provider (EnvironmentSecretProvider, DictionarySecretProvider) - StringResolver for inline expressions in YAML ({$.field},
$secrets.x, $ {ENV_VAR}) - JsonFileSecretProvider — local development secrets from
secrets.json - AppConfigurationSecretProvider — Azure App Configuration with label-based environments
- CompositeSecretProvider — chained resolution: secrets.json → App Configuration → env vars
- Key Vault references — App Configuration stores pointers to Azure Key Vault secrets instead of raw values. Key Vault provides fine-grained per-key access policies, enabling secret isolation between pipeline maintainers (e.g., maintainer A manages
crm-newsletter/*, maintainer B managesproduct-sync/*). Elwood's ISecretProvider resolves Key Vault references transparently. - Per-pipeline secret scopes — each pipeline declares a
secrets.scopein YAML (e.g.,crm-newsletter).$secrets.passwordresolves as{scope}/passwordin the backing store. The runtime enforces the scope — a pipeline cannot read another pipeline's secrets. Cloud-agnostic (works with any store that supports key prefixes). - Full destination type schema (11 types — schema defined in docs, code deferred to Step 4)
Step 2 — CLI Executor: ✅
-
elwood pipeline run <yaml> --source name=filecommand -
--source-envelope name=filefor explicit envelope files (no auto-detection) -
--output-dirwrites each output as{name}.json -
elwood pipeline validate <yaml>— validates YAML, scripts, dependencies, duplicates - 5 pipeline test scenarios (single source, multi-source merge, XML→CSV, fan-out, depends chain)
- Pipeline conformance test runner (discovers spec/pipelines/*)
- 21 CLI integration tests (including 6 pipeline tests)
Step 3 — State + persistence: ✅
- ExecutionState model: per-execution, per-source, per-output step state with timestamps
-
IStateStore+IDocumentStoreinterfaces -
InMemoryStateStore+InMemoryDocumentStoreimplementations -
FileSystemStateStore+FileSystemDocumentStorefor persistent local state - PipelineExecutor tracks state automatically when stores are provided
-
elwood pipeline status [state-dir]— shows recent executions with step details -
--output-dirpersists state to.state/subdirectory
Step 4 — Sync Executor:
-
ISourceConnector+IDestinationConnectorinterfaces -
HttpSourceConnector— fetch from REST APIs -
FileSourceConnector— read from local/network paths -
HttpDestinationConnector— deliver to REST APIs -
FileDestinationConnector— write to local/network paths -
SyncExecutor— end-to-end execution: trigger + pull sources, maps, IDM, outputs, destinations - 7 SyncExecutor tests with mock HTTP (no real network calls)
-
elwood pipeline serve <yaml>— start HTTP listener for trigger sources (deferred) - POST body from script —
HttpSourceConnectorsupportsbodyfield onfrom.httpconfig: an inline expression or.elwoodscript reference evaluated against the IDM, serialized as the POST/PUT request body - Accepted status codes —
HttpSourceConnectorsupportsacceptedStatusCodesfield (e.g.,"2xx,4xx,5xx"): captures the response on non-2xx instead of throwing. Status code exposed via$source.http.statusCodein source map scripts - Dynamic response status code —
OutputConfigsupportsresponseStatusCodefield: expression evaluated against the IDM that sets the HTTP response status code (e.g.,$.crmStatusCode).PipelineResultcarries the resolved status code for the HTTP trigger to use - HTTP trigger auth — optional
authsection on HTTP trigger sources (type: basic,user/passwordfrom$secrets). Runtime validates the Authorization header; omit for no auth
Step 5 — Deployment + Runtime API:
-
IPipelineStoreinterface — source of truth for pipeline YAMLs + .elwood scripts-
FileSystemPipelineStore— local folder (dev/CLI) -
GitPipelineStore— git repo as backing store (deferred to Step 6)- Every save = git commit (automatic versioning, diff, audit trail)
- Revisions API =
git log, restore =git checkout+ commit - Deploy = tag or push to deploy branch
- Backed by any git remote (Azure DevOps, GitHub, GitLab, local bare repo)
- Developers can edit in VS Code and push — portal is optional
- Each pipeline is a folder:
{pipeline-id}/pipeline.elwood.yaml+{pipeline-id}/*.elwood
-
-
IPipelineRegistryinterface defined — Redis impl deferred to Step 6- Route table: endpoint patterns → pipeline ID (for HTTP request matching)
- Pipeline content cache: full YAML + all .elwood scripts stored in Redis (~30KB per pipeline)
- Search index: pipeline names + content for portal search
- Webhook-triggered sync: git push → API server pulls → reads changed files → updates Redis
- Incremental updates for normal commits, full rebuild on startup
- Executors are fully stateless — read everything from Redis, no local git clone needed
-
elwood deploycommand — writes to IPipelineStore (git) + updates IPipelineRegistry (Redis) -
Elwood.Runtime.Api— REST API layer (ASP.NET minimal API project created) - API reads/writes pipelines via
IPipelineStore(Redis registry deferred) - Pipelines:
-
GET /api/pipelines— list pipelines, filter by name -
POST /api/pipelines— create new pipeline -
GET /api/pipelines/{id}— get pipeline YAML + associated .elwood scripts -
PUT /api/pipelines/{id}— update pipeline -
DELETE /api/pipelines/{id}— delete pipeline -
GET /api/pipelines/{id}/revisions— version history (returns empty for FileSystem store) -
POST /api/pipelines/{id}/revisions/{rev}/restore— restore to previous version (deferred with Git store) -
POST /api/pipelines/{id}/validate— runelwood validate -
POST /api/pipelines/{id}/deploy— deploy to pipeline store (deferred)
-
- Scripts (
.elwoodfiles associated with a pipeline):-
GET /api/pipelines/{id}/scripts/{name}— get script content -
PUT /api/pipelines/{id}/scripts/{name}— create or update a script -
DELETE /api/pipelines/{id}/scripts/{name}— delete a script -
POST /api/pipelines/{id}/scripts/{name}/test— run script against provided input, return result
-
- Executions:
-
GET /api/executions— list executions, filter by pipeline/status/time range -
GET /api/executions/{id}— full execution state from IStateStore -
POST /api/executions— trigger a pipeline run -
DELETE /api/executions/{id}— cancel a running execution (deferred)
-
- Documents:
-
GET /api/documents/{ref}— retrieve payload/output from IDocumentStore (deferred)
-
- System:
-
GET /api/health— runtime health status -
GET /api/metrics— running executions count, recent activity summary
-
- Auth: JWT bearer tokens (MSAL / Azure AD integration) — deferred to portal phase
Step 6 — Cloud Executors (separate packages):
Broken into seven sub-steps. The cloud runtime supports two execution modes via a mode: field in pipeline YAML — sync (HTTP function runs end-to-end and returns one designated output) and async (HTTP function returns 202 + executionId, fan-out via Service Bus).
Step 6a — Schema additions + Azure storage adapters: ✅
-
mode: sync | asyncin pipeline YAML (default: sync) -
response: trueon outputs — sync mode requires exactly one, async forbids - Validation in
PipelineParser.ValidateConfig()— wired intoParse() -
Elwood.Pipeline.Azureopt-in NuGet package (net8.0;net10.0)-
RedisStateStore— Lua-script-atomic per-step updates, KEEPTTL preserved, 3-day default TTL -
BlobDocumentStore— Azure Blob with auto-create container, lifecycle delegated to blob lifecycle policies -
RedisPipelineRegistry— read-only and writable constructors, content cache + literal route matching + search -
AddElwoodAzureStorage(...)DI helper + writable variant for the API server
-
- 24 integration tests via Testcontainers (Redis 7-alpine + Azurite)
- Concurrent-writers test proves Lua atomicity (20 parallel updates, no lost data)
- KEEPTTL test proves TTL preservation across updates
- CI fix —
dotnet testnow runs the entire solution (was: only Core.Tests) - All 6 .NET packages + 2 npm packages bumped to 0.4.0
Step 6b — GitPipelineStore: ✅
-
GitPipelineStorewraps the git CLI — every save = git commit, GetRevisions = git log, RestoreRevision = git show + commit -
GitHelperutility: thin wrapper around the git CLI (no LibGit2Sharp — avoids native binary issues on newer .NET) - Backed by any git remote (Azure DevOps, GitHub, GitLab, local bare repo) — the store manages commits, the API server manages push/pull
- 11 tests: CRUD, revision history with limits, restore (including script add/removal), author tracking, empty repo safety
Step 6c — AsyncExecutor (step-at-a-time, queue-driven): ✅
-
AsyncExecutor— step-at-a-time execution engine formode: asyncpipelinesStartAsync(pipeline, payload)— creates state, stores trigger payload + pipeline content in IDocumentStore, queues stage 0 sourcesExecuteStepAsync(message)— processes one source or output step per invocation
-
IStepQueueinterface +InMemoryStepQueuefor tests (Service Bus impl ships in 6d) -
StepMessagemodel (ExecutionId, PipelineId, StepType, StepName, StageIndex) - Fan-in via idempotent steps: after completing a source, checks all sources in stage; duplicates are no-ops
- Stage plan stored in IDocumentStore so queue workers can reload independently
- Pipeline content stored in IDocumentStore — workers are stateless, no local git clone
- 8 tests: start + state creation, source processing, output processing, multi-stage ordering, concurrent sources, idempotency, failure handling, end-to-end
Step 6d — Elwood.Runtime.Azure Functions project:
-
BlobPipelineStore— implementsIPipelineStoreover Azure Blob Storage (folder-per-pipeline layout in a blob container). ReplacesFileSystemPipelineStorefor serverless environments where there is no persistent disk. - HTTP trigger (catch-all): matches route via
RedisPipelineRegistry, dispatches toSyncExecutor(sync) or starts execution + queues first step (async) - Service Bus queue trigger: invokes
AsyncExecutor.ExecuteStepAsync(...)(async mode only — not needed for sync-only pipelines) - DI wiring:
AddElwoodAzureStorage(...)+ Functions startup - HTTP method support added to
SourceConfigschema - Route pattern matching with parameter extraction (e.g.
/api/{category})
Step 6e — Pipeline deployment flow:
How pipelines move from editing to live execution. Two modes, detailed in docs/pipeline-deployment-modes.md:
- Mode 1 (save = live): Portal saves → API writes to Blob Storage → updates Redis → pipeline is live immediately. Git commit as optional audit trail. Good for solo developers and dev/test environments.
- Mode 2 (save = draft, merge = live): Portal saves to a git branch → submits PR → review + merge → webhook triggers API → updates Blob + Redis → pipeline goes live. Good for teams and production. Requires: portal branch awareness, PR creation, merge webhook handler.
-
elwood deploy <yaml>CLI command — uploads pipeline to blob + updates Redis - Git commit on save (Mode 1 audit trail) — optional, configured per environment
- Merge webhook handler (Mode 2) — receives git merge events, updates blob + Redis
Step 6f — Terraform:
-
infra/azure/main.tf— Storage Account (pipelines container + Functions runtime) + Redis + Function App - Support reusing existing resources via
datablocks (resource group, App Service Plan, App Configuration, Application Insights) -
infra/azure/variables.tf— resource names, region, SKU, existing resource references -
infra/azure/outputs.tf— function URL, storage connection, Redis connection -
infra/azure/terraform.tfvars.example— sample values (real tfvars gitignored) - Getting started documentation for
terraform apply→ deploy → test
Step 6g — End-to-end integration test:
- docker-compose: Azurite + Redis + Service Bus emulator
- Run a real pipeline through the Functions runtime locally
- Verify state, documents, fan-out, route matching all work end-to-end
Step 6 — AWS Executor (later):
- Lambda + SQS + DynamoDB + S3 adapters (separate package:
Elwood.Pipeline.Aws) - AWS executor Functions equivalent (Lambda functions package)
Infrastructure (separate repo: elwood-infra):
- Terraform module: Azure (Function App + ASB + Storage + App Insights)
- Terraform module: AWS (Lambda + SQS + DynamoDB + S3)
- Example configurations (minimal, production)
Goal: A web-based management UI for authoring, deploying, testing, and monitoring Elwood integration pipelines. Re-engineered from an existing enterprise integration frontend, adapted for Elwood's pipeline YAML + script architecture.
Tech stack: Next.js + React + Tailwind + Monaco Editor + Redux (separate repo: elwood-portal)
Implementation prompt: See docs/prompts/implement-portal.md
- Browse, search, create, edit pipeline YAML files (
.elwood.yaml) - Browse, search, create, edit Elwood scripts (
.elwood) with full Monaco syntax highlighting + autocomplete - Version history with diff view and restore (via GitPipelineStore)
- Deploy pipelines to runtime
- Bulk upload (ZIP of pipeline + scripts)
- Validation:
elwood validateintegrated in the editor
- Run
.elwoodscripts against input data with live preview - Load payloads from execution history or file upload
- Format-aware input: JSON, CSV, XML, Text with conversion preview
- Compare input vs output side-by-side
- Pipeline execution dashboard (reads from
IStateStore) - Real-time activity log with filtering (by pipeline, status, time range)
- Execution detail view: steps, fan-out progress, errors, duration
- Correlation tracing across multi-step async flows
- Health status indicator for the runtime
- Browse execution state (reads from
IStateStore) - Browse stored documents / IDM (reads from
IDocumentStore) - Download payloads and outputs
- Role-based access control (Admin, Editor, Viewer)
- Authentication (Azure AD / MSAL)
- Multi-environment support (dev, staging, production)
- Integrated Playground — run Elwood expressions in the browser via
@elwood-lang/corenpm package - 86 official examples bundled from
spec/test-cases/(static JSON, build-time) - Example sidebar with search, categories, click-to-load
- API-backed examples (refinement): replace static bundle with
GET /api/examplesendpoint. Official examples embedded as assembly resources inElwood.Pipeline(read-only, updated on NuGet upgrade). Organization examples stored server-side viaPOST/PUT/DELETE /api/examples/{id}(same CRUD pattern as pipelines). Portal shows two sections: "Elwood" (official) and "Organization" (user-managed). Enables each adopting organization to maintain their own shared example library alongside the official ones.
- Project setup (Next.js + Tailwind + Monaco)
- API layer consuming
Elwood.Runtime.Apiendpoints (real API, not mocks) - Pipeline list page with search
- Pipeline editor with Monaco + Elwood syntax highlighting + script tabs
- New pipeline creation with uniqueness check (409 Conflict)
- Core pages: execution dashboard, execution detail
- Testing/preview panel (run pipeline against input from the editor)
- Version history with diff view (connects to GitPipelineStore revisions API)
- Authentication + role-based access
- Deployment integration
Goal: IDE support, developer tools, and community ecosystem.
- VS Code extension: syntax highlighting for
.elwoodand.elwood.yamlfiles - Language server (LSP): autocomplete, go-to-definition for
letbindings, error squiggles - Hover documentation for built-in methods
- Snippet library for common patterns
- Playground multi-format expansion: input/output format selectors, conversion preview, dual-view output (see playground-spec.md Phase 2+ section)
- Trace mode:
elwood run --traceshows step-by-step pipeline execution with intermediate values - Visual debugger: step through pipe stages, inspect values at each step
- Schema inference: analyze an Elwood expression and infer the expected input/output JSON schema
- Documentation site (GitHub Pages or similar)
- Example library: real-world transformation patterns
- Community contributions: custom function plugins
Goal: A JSON database where you store documents and query them with Elwood. No predefined schema, no document size limits. PostgreSQL as the storage backend.
Detailed design: See docs/architecture-vision.md Phase 5 Vision section.
- Elwood as the native query language — the same language for querying and transforming, no context switch
- No document size limits — automatic splitting handles 100MB+ documents
- Schema-on-write — define structure when storing, not upfront
- PostgreSQL backend — we build a query translation layer, not a database engine
- Fixed schema: 2 tables (
collections+chunks), no dynamic table creation - Large documents are split along a configured path (e.g.,
$.orders[*]→ one row per order) - PostgreSQL JSONB indexing (GIN + expression indexes) handles query performance
- Elwood queries translate to SQL, with non-SQL operations (
.toUpper(), complexselect) executed in Elwood on the result set
-
Elwood.Dbproject (separate repo:elwood-db) - PostgreSQL schema (collections + chunks tables)
- Collection management (
elwood db create,elwood db drop) - Document storage with automatic splitting
- Index management (translate user-defined paths to PostgreSQL expression indexes)
- Query translation: Elwood
where→ SQLWHERE,take/skip→LIMIT/OFFSET,orderBy→ORDER BY - Aggregation push-down (
count,sum,min,max→ SQL) - CLI integration:
elwood db query,elwood db store - REPL integration:
:db connect,:db use - Optional SQLite backend for embedded/dev use
Phase 1 ✅ Build the .NET engine
↓
Phase 1b ✅ Lazy evaluation (.NET) → repo restructure → TypeScript port
↓
Phase 1c ✅ Publish everything (GitHub, NuGet, npm, CI, Playground, API container)
↓
Phase 2 ✅ Multi-format I/O (fromCsv, toXml, etc.) + script-based maps
↓
Phase 2b ✅ Performance — 2x faster than legacy baseline (compiled mode explored, not needed)
↓
Phase 3 Integration pipeline configuration (Elwood Runtime + Executors)
↓
Phase 3b Elwood Management Portal (web UI for authoring, testing, monitoring)
↓
Phase 4 IDE support, developer tools, ecosystem
↓
Phase 5 Elwood DB — JSON database with Elwood queries (separate repo)
Phase 3b is a separate repo (elwood-portal). Phase 5 is a separate repo (elwood-db).
Elwood is designed for incremental adoption:
- Phase 1 (done): Standalone transformation engine. Use via CLI or .NET library.
- Phase 1b (done): Cross-platform reach. Use in browsers, Node.js, edge runtimes.
- Phase 1c (done): Open-source launch. Available via NuGet, npm, browser playground, and self-hosted API container.
- Phase 2 (done): Multi-format I/O. All formats complete. Extension API for XLSX. CLI format flags.
- Phase 2b (done): Performance verified — 2x faster than legacy baseline on 100K rows. Compiled mode explored but interpreter is already optimal.
- Phase 3: Integration pipelines. YAML defines sources, transforms, and destinations. Pluggable executors run them.
- Phase 3b: Management portal. Web UI for authoring pipelines, testing transformations, monitoring executions.
- Phase 4: IDE support, developer tools, and community ecosystem.
- Phase 5: Elwood DB. Store JSON, query with Elwood. PostgreSQL backend, no size limits.
Each phase is independently useful. No phase requires adopting a later one.
{elwood}
As a library → NuGet: Elwood.Core / Elwood.Json
npm: @elwood-lang/core
As a CLI tool → Download from GitHub Releases (Windows, macOS, Linux)
Or: dotnet tool install --global Elwood.Cli
In the browser → Playground: https://max-favilli.github.io/elwood/
As an API → docker run -p 8080:8080 ghcr.io/max-favilli/elwood-api
POST /api/evaluate { script, input } → result
(integrates with any iPaaS via HTTP)