Skip to content

feat(dmap): add distributed map operation - #2

Draft
nvasiu wants to merge 2 commits into
mainfrom
feat/map-run
Draft

nvasiu wants to merge 2 commits into
mainfrom
feat/map-run

Conversation

@nvasiu

@nvasiu nvasiu commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

Summary

Adds the distributed map operation (ctx.distributed_map) to the Python SDK:
A map run processes a bounded dataset in parallel. A customer starts a map run from a durable function, naming a source to read items from, a processor function to invoke per batch, and concurrency, retry, and failure settings. The service reads items from the source, groups them into batches, invokes the processor for each batch, retries failures, tracks progress, routes successful results and failed items to destinations, and reports completion.

Changes

dmap/models.py

The result types a customer receives back from a map run.

  • DistributedMapSummary: what ctx.distributed_map returns, describes the run's overall outcome.
  • DistributedMapResult: returned when a DistributedMapResultConfig is passed. Contains individual map run item outcomes.
  • DistributedMapResultItem and DistributedMapItemError: represent a single item's result / error.

dmap/__init__.py

Empty package marker for the dmap package.

config.py

The input types a customer constructs to describe a map run, and the distributed map enums.

  • DistributedMapConfig: optional settings for distributed map.
  • DistributedMapResultConfig: subclass that additionally collects item results inline, selecting the return type statically.
  • InlineSource, S3Source, ReaderSource: describe where map run items come from.
  • DistributedMapProcessor: describes the Lambda that processes items, how outcomes are reported back, and how failing items are retried via max_retry_attempts and max_retry_duration.
  • DistributedMapCompletionConfig: defines item failure thresholds for marking the overall map run failed.
  • S3Destination, DistributedMapOnSuccessConfig, DistributedMapOnFailureConfig, DistributedMapDestinationConfig: for routing successful and failed item records to S3.
  • The enums live here rather than in lambda_service, so config.py has no runtime import from it.

context.py

The customer-facing entry point on the durable execution context.

  • ctx.distributed_map: the method a customer calls to run a distributed map. Overloaded so passing a DistributedMapResultConfig types the return as DistributedMapResult and anything else as DistributedMapSummary.
  • Validates max_concurrency against the documented ceiling of 10000.

dmap/handlers.py

Authoring decorators for the processor Lambda, so a customer writes a plain function rather than the item or batch protocol.

  • distributed_map_item_handler, distributed_map_batch_handler, distributed_map_reader, and the durable variants durable_distributed_map_item_handler and durable_distributed_map_batch_handler.
  • Named and shaped like durable_execution, so they work bare, with options, or called directly with a function.
  • Responses are built from envelope dataclasses rather than hand-built dicts.

operation/dmap.py

The executor that drives the operation against the durable execution runtime.

  • Suspends the caller's function while the run executes and resumes it with the finished outcome.
  • Every terminal state resolves, and throw_if_error() opts into raising.
  • All translation between the config types and the service shapes happens here.

lambda_service.py

Serialization for carrying the operation and its results to and from the backend service.

  • One dataclass per API shape, named after the shape rather than with a Wire suffix.
  • Items, Output, and the record body are opaque strings, matching the API model, so any serdes works rather than JSON only.

state.py

Durable execution state handling, so a run's outcome persists across suspend and resume.

  • Records that a distributed map operation carries its outcome on the operation itself rather than as an operation-level result or error.

exceptions.py

The error type a customer catches.

  • DistributedMapError: raised when a run or an item fails.

plugin.py

Operation type registration.

  • Adds DISTRIBUTED_MAP to OperationType.

__init__.py

The package's public API surface.

  • Exports the result types, their enums, and the five authoring decorators, matching how durable_execution and durable_step are exported.
  • The input config and factory types are imported from config.

Tests

Tests are split by layer so each sits alongside its module.

tests/config_test.py

  • Config and argument validation.

tests/lambda_service_test.py
tests/dmap/models_test.py

  • Serialization round trips and the result types.

tests/operation/dmap_test.py
tests/context_test.py

  • Executor behaviour and the ctx.distributed_map surface.

tests/dmap/handlers_test.py

  • Authoring wrapper tests: checking that they process items, report failures, reject bad inputs.

tests/e2e/dmap_int_test.py
tests/e2e/dmap_helpers_int_test.py

  • End to end tests mocking backend responses: suspend / resume, collect results, throw on failure.

Future Tasks

  • Add distributed map to the local emulator (in the testing package).
    • When this is done, we can add full end to end tests using the emulator.
  • Add distributed map examples to the examples package.

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.

@nvasiu
nvasiu force-pushed the feat/map-run branch 6 times, most recently from cfa93bd to cbacaaf Compare September 23, 2026 17:56
@nvasiu

nvasiu commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner Author

@yaythomas

I accidentally broke the last PR. So here's a new clean PR to add the distributed map operation with the restructure you suggested in your comment here (along with fixes to other issues you pointed out).

But I made the these modifications to your proposed restructure:

  • Your suggestion: All enums should be defined in lambda_service, and customer visible enums should be imported into config.py.

    • My change: I moved customer facing enums into config, and kept the rest of the enums in lambda_service. So lambda_service imports what it needs from config, but config does not import from lambda_service.
    • Justifications:
      • Avoid having config importing from the serialization layer. Follows the contributing guide's rule that config should be the lowest level import.
      • Some enums (like DistributedMapStatus) aren't used in the serialization layer, so it doesn't make sense to keep them in lambda_service.
  • Your suggestion: Export nothing new from the package root.

    • My change: I export the following from the package root: result types, customer facing enums, error type and the authoring helpers.
    • Justifications:
      • This change is consistent with other existing operations. Other operation's result types, customer facing enums and durable_* decorators are exported from the root.
      • Exporting these structures from the root lets customers import them without needing to reach into submodules.

I also used dmap for directory and file names, rather than distributed_map. You suggested dmap in an older comment to match naming for existing files / avoid underscores.

Thoughts on these changes?

I made this graph of the new operation structure to more easily see how the components relate:

image

Besides the restructure, I also made the following fixes based on your comments:

  • Terminal failures resolve instead of raising.
  • Items, Output and the record body pass through the serdes as strings, so non-JSON serdes work.
  • Removed distributed_map_id. This just returned the last segment of the map run ARN, so it wasn't that useful. We decided to remove this from all SDKs.
  • Factories renamed and flattened (batch, item_failures, item_results, InlineSource, S3Source, ReaderSource, S3Destination).
  • Removed union return type. Now the return type is selected statically by config type.
  • Translation has all been moved to the executor.
  • Renamed serialization dataclasses after their API shapes, like existing ones.
  • Envelope dataclasses to replace any hand built dicts.
  • Removed the module global _CTX = SerDesContext(). Now context is built per call.
  • Authoring helpers were turned into decorators and renamed to match existing decorators.
  • Removed the unrelated execution.py invoke change.

@nvasiu
nvasiu force-pushed the feat/map-run branch 2 times, most recently from ee54c16 to 50f5eba Compare September 25, 2026 17:53
Add ctx.distributed_map with inline, S3, and reader sources, the config,
processor, completion, and destination types, the result types, and
function-authoring helpers for item and batch handlers.

Serialization lives in lambda_service.py with one dataclass per API shape.
Translation happens only in the executor, so config.py has no runtime
import from lambda_service. Tests are split by layer.

Items, Output, and the record body are opaque strings, matching the API
model, so any serdes works rather than JSON only.

Every terminal state resolves with the summary, and throw_if_error() opts
into raising. Missing details or a missing completion reason raise
ExecutionError.

The processor factories are batch, item_failures, and item_results. Source
and destination factories are InlineSource, S3Source, ReaderSource, and
S3Destination. Retry is max_retry_attempts and max_retry_duration on the
processor. Passing DistributedMapResultConfig selects the DistributedMapResult
return type statically. Only the result types and their enums are exported
from the package root.

@yaythomas yaythomas left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great rewrite, thank you very much! Status contract, string-vs-JSON on the wire, factory names, package-root exports, the return-type union. The envelope dataclasses in dmap/handlers.py and the DistributedMapResultConfig overload are especially nice. Thank you for the second pass.

Headlines:

  • Please bring back a distributed_map_id accessor, derived properly this time.
  • The module-level _build_* helpers in operation/dmap.py want to be from_* factories on the wire dataclasses (pure shape translation) or executor methods (anything that serializes). That also resolves the remaining Any parameters.
  • list[T] rather than tuple[T, ...] for the collection fields.

Details inline.

unprocessed_count: int
distributed_map_run_arn: str | None = None
completion_details: str | None = None
total_count: int | None = None

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please could you add back a distributed_map_id derived from the ARN? The last round flagged the buggy arn.rsplit(":", 1)[-1]; the right fix is to derive it correctly rather than drop the accessor, since the run id is the handle customers need when locating a run's output objects.

Rather than a bare split, mirror DurableExecutionArn.from_arn in types.py: a frozen dataclass with an anchored regex, returning None when the ARN doesn't match. Why? A split returns something for any string, including the wrong segment for an unexpected shape; the regex validates the whole grammar and gives the parsing one home and one test.

_DISTRIBUTED_MAP_RUN_ARN_PATTERN = re.compile(
    r"^(arn:[^:]*:lambda:[^:]*:[^:]*:function:[^:/]+:[^:/]+"
    r"/durable-execution/[^/]+/[^/]+)/distributed-map-run/([a-z0-9]+)$"
)

@dataclass(frozen=True)
class DistributedMapRunArn:
    durable_execution_arn: str
    run_id: str

    @classmethod
    def from_arn(cls, arn: str) -> DistributedMapRunArn | None:
        match = _DISTRIBUTED_MAP_RUN_ARN_PATTERN.match(arn)
        if not match:
            return None
        return cls(durable_execution_arn=match.group(1), run_id=match.group(2))

Then distributed_map_id is a property that calls from_arn and returns parsed.run_id if parsed else None.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Will follow this approach, but I'll reuse the ARN regex in types.py to validate, so we don't have the same regex in 2 places.

}


def _build_inline_items(

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The translation helpers are module-level functions and they mix two kinds of work. Splitting them along that line would match the rest of the codebase:

  1. Pure shape translation (customer config in, wire dataclass out, no serdes, no execution context): _build_s3_source_config, _transform_for, _build_processor_config, _build_completion_config, _destination_entry_fields, _build_destination_config. Please make these from_* classmethods on the target lambda_service dataclass: DistributedMapS3SourceConfig.from_source(s3: S3SourceConfig), DistributedMapProcessorConfig.from_processor(processor), DistributedMapOnSuccessConfig.from_destination(d: SuccessDestination), and so on. Why? CONTRIBUTING: "Encapsulate conversion logic in a from_x factory and to_x method on a class"; ErrorObject.from_exception is the existing example of a wire class building itself from a domain object. lambda_service importing config is the right direction (config is the lowest layer), so no cycle. _UNLIMITED_RETRY_WIRE and _RESPONSE_TYPE_FOR_MODE move next to the processor class. The kwargs-dict splat in _destination_entry_fields disappears once each subclass builds itself, and each factory takes its concrete config type, which retires the Any parameters (CONTRIBUTING §Typing). DistributedMapCompletionConfig exists in both modules; alias one at the import site, as concurrency/models.py does with BatchResult as BatchResultProtocol.

  2. Serialization (_build_inline_items, the reader initial_state branch of _build_source_config): these call serialize() with the operation id and ARN, which only the executor has, and lambda_service can't import serdes without a cycle, so they stay in the executor. Please make them methods though: today each takes operation_id and durable_execution_arn as parameters the caller reads off self, and _build_distributed_map_options takes six parameters, all six of which are self.*, from one call site. A small self._serialize(serdes, value) reads the two context values once. Then the wire class takes strings: DistributedMapSourceConfig.create_inline(items: list[str], max_items=), .create_s3(s3, max_items=), .create_reader(function_name, initial_state: str | None, max_items=), in the style of OperationUpdate.create_step_start / create_invoke_start. InvokeOperationExecutor.check_result_status is the existing example: serialize with self.* context, hand the string to the wire constructor. (serialize() already defaults serdes=None to the JSON serdes, so the or DEFAULT_JSON_SERDES can go too.)

  3. Result reconstruction (_resolve_summary, _distributed_map_status_from_operation): consider DistributedMapSummary.from_operation(operation) in dmap/models.py, with CheckpointedResult.create_from_operation as the precedent. DistributedMapResult additionally needs the deserialized items, which stay an executor concern, so DistributedMapResult.from_operation(operation, items).

After this operation/dmap.py is the executor class and nothing else, which is what operation/invoke.py looks like. Mostly moving code rather than writing it.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the suggestions, preparing an update now with a couple changes:

serialize() already defaults serdes=None to the JSON serdes, so the or DEFAULT_JSON_SERDES can go too.

serialize() currently defaults to EXTENDED_TYPES_SERDES, so removing DEFAULT_JSON_SERDES changed the result and failed the tests. I'll need to keep that.

DistributedMapResult additionally needs the deserialized items, which stay an executor concern, so DistributedMapResult.from_operation(operation, items).

DistributedMapResult subclasses DistributedMapSummary which already has from_operation(operation). So if I try to override it with an extra required parameter (from_operation(operation, items)), mypy will reject it.

Instead I'll give DistributedMapResult a new from_operation_and_items(operation, items). And I will overload DistributedMapResult.from_operation() to return an error saying its not valid on a result, and to use from_operation_and_items() instead.

class DistributedMapInlineSourceConfig:
"""Represent the items an inline map run source reads from."""

items: tuple[str, ...] = ()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

list[MyType] is probably more natural here than tuple[MyType, ...], and throughout the codebase it's mostly list[T] for this sort of thing (BatchResult.all, CheckpointUpdatedExecutionState.operations, ErrorObject.stack_trace). It's true tuple gives an immutability that could arguably belong on a frozen dataclass, but Sequence[T] / list[T] is the more common shape for a single-item-type collection. Same for the other tuple fields in config.py and dmap/handlers.py; the tuple(...) / list(self.x) conversions go with them. Pls remember to keep the copy in InlineSource.of (list(items)) so a caller mutating their list afterwards doesn't change what gets checkpointed.


PASS_THROUGH_SERDES: SerDes[Any] = PassThroughSerDes()

_MAX_CONCURRENCY_LIMIT = 10000

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

_MAX_CONCURRENCY_LIMIT = 10000 lives in context.py, next to nothing else about distributed map. Consider a ClassVar on DistributedMapProcessor beside UNLIMITED, or next to the wire dataclass that carries the value, so the bound sits with the thing it bounds. Trivial.

# Replay with the run completed.
_map_run_id, replay_event = _replay_event(
{
"Status": "SUCCEEDED",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The DistributedMapDetails fixtures carry a "Status": ... key, but DistributedMapDetails.from_dict doesn't read one (status comes from the operation, which the fixture also sets). Harmless since the key is ignored, but it suggests a shape the code doesn't have; dropping it would make the fixtures match from_dict.

raise ValidationError(msg)


def distributed_map_item_handler(

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The concurrency default moved from len(records) to 1, and the docstring says so. I think that's the right default: the old one required customer code to be thread-safe out of the box, which is a surprising precondition. Just recording the deliberate behaviour change, no action needed.

summary = self._resolve_summary(operation)
return CheckResult.create_completed(summary)

# Operation-level terminal failure. Every terminal state resolves with the

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This branch is exactly right: every terminal status resolves with the summary, and ExecutionError is raised only when the backend omitted the details block. With test_terminal_failure_with_details_resolves_with_summary and test_operation_level_terminal_failure_without_details_raises the contract is well pinned. Nice!

)
return executor.process()

@overload

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cleanly done. config: DistributedMapResultConfig selects DistributedMapResult, everything else the summary, no isinstance at the call site. Nice!

several threads.
"""
if func is None:
return functools.partial(

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if func is None: return functools.partial(...) mirrors durable_execution, so bare, with-options and direct-call forms all work, and the envelope dataclasses mean no dict[str, Any] plumbing anywhere. Nice!

@nvasiu

nvasiu commented Oct 2, 2026

Copy link
Copy Markdown
Owner Author

@yaythomas
Another update to address the above comments, and also some misc. changes to match repo conventions better.

Summary

Adding back distributed_map_id:

  • dmap/models.py: added DistributedMapRunArn with a from_arn classmethod.
  • dmap/models.py: added the distributed_map_id property to DistributedMapSummary.

Moving translation helpers:

  • operation/dmap.py: deleted the eleven module-level translation helpers.
  • lambda_service.py: added a from_* classmethod to each wire type that translates one config-layer counterpart.
  • lambda_service.py: added _transform_for to derive the S3 transform.
  • lambda_service.py: added a create_* classmethod to DistributedMapSourceConfig for each of its three source variants.
  • lambda_service.py: added _create to the base the two destination entries share, replacing the dict-unpacking construction.
  • lambda_service.py: moved _UNLIMITED_RETRY_WIRE and _RESPONSE_TYPE_FOR_MODE in from operation/dmap.py.
  • lambda_service.py: imported the config-layer completion config under the alias ConfigCompletionConfig.
  • operation/dmap.py: added the _serialize, _serialize_inline_items, _reader_initial_state, _build_source_config and _build_options methods, reading the operation id and execution ARN off self.
  • operation/dmap.py: dropped the or DEFAULT_JSON_SERDES from each call site and applied it inside _serialize.
  • operation/dmap.py: typed the serdes parameter of _deserialize_items as SerDes | None.
  • dmap/models.py: added DistributedMapSummary.from_operation and DistributedMapResult.from_operation_and_items.
  • dmap/models.py: added DistributedMapResult.from_operation, which raises ValidationError naming the factory that takes the items.
  • dmap/models.py: moved _resolve_terminal and the status mapping in from operation/dmap.py.

Using list instead of tuple:

  • config.py, lambda_service.py, dmap/handlers.py: retyped the collection fields and the _validate_columns parameter to list.
  • lambda_service.py, dmap/handlers.py: removed the list(...) copies from to_dict.

Moved _MAX_CONCURRENCY_LIMIT:

  • lambda_service.py: moved _MAX_CONCURRENCY_LIMIT in from context.py, above DistributedMapOptions.
  • context.py: imported the constant from lambda_service.

Removed status from e2e fixtures:

  • tests/e2e/dmap_int_test.py: removed the Status key from the three DistributedMapDetails fixtures.

Other changes:

  • config.py: moved the nine source factories onto DistributedMapSource as classmethods, named for the source kind, and deleted the InlineSource, S3Source and ReaderSource factory classes.
  • config.py: moved the two destination factories onto the destinations as from_uri, and deleted S3Destination.
  • config.py: dropped the Config suffix from the three resolved types that are not a config= argument, moving the destination holder below the two destinations it holds.
  • dmap/handlers.py: replaced the report string argument on both item handler decorators with response_mode, typed ProcessorResponseMode.
  • dmap/handlers.py: renamed _validate_report to _validate_response_mode and made it reject BATCH.
  • dmap/handlers.py: changed the BatchItemStatus import to config.
  • __init__.py: exported ProcessorResponseMode.
  • dmap/models.py: renamed the status mapper to _to_distributed_map_status.
  • operation/dmap.py: replaced the direct checkpoint fetch with self._get_checkpoint_result().
  • tests/operation/dmap_test.py: added sub_type and name to 24 operation fixtures.
  • tests/e2e/dmap_int_test.py: made the checkpoint fake echo sub_type and name back from the update.
  • tests/operation/dmap_test.py: added test_checkpoint_from_a_different_map_is_rejected.
  • tests/dmap/handlers_test.py: moved three tests and their helper out to tests/dmap/models_test.py and tests/operation/dmap_test.py, and removed the twelve imports that went unused.
  • tests/dmap/models_test.py: added tests for the run ARN, distributed_map_id, and the from_operation rejection.
  • tests/config_test.py, tests/operation/dmap_test.py, tests/e2e/dmap_helpers_int_test.py: updated the call sites for the renamed factories and the response_mode argument.

- The wire types translate their config-layer counterparts through
  from_* and create_* classmethods, so operation/dmap.py holds the
  executor alone and its serialization reads the operation id and
  execution ARN off self.
- dmap/models.py rebuilds the summary and the result from a terminal
  operation, and parses the run ARN to expose the run id.
- The source and destination factories are classmethods on the type
  they return, and the resolved types that are not a config= argument
  drop the Config suffix.
- Collection fields are list rather than tuple, matching the rest of
  the SDK.
- The item handler decorators take a typed response_mode in place of
  a report string.
- The concurrency bound and the wire sentinels sit beside the types
  they apply to, and each shared type is imported from the module
  that defines it.
- The executor fetches its checkpoint through the base class helper,
  which raises NonDeterministicExecutionError on an identity
  mismatch.
- Each test file covers the module it is named for.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants