Skip to content

Fix time range filters, non-ASCII search and oversized values in the EF Core persisters - #5934

Merged
johnsimons merged 2 commits into
masterfrom
john/ef_primary_fixes
Sep 30, 2026
Merged

johnsimons merged 2 commits into
masterfrom
john/ef_primary_fixes

Conversation

@johnsimons

@johnsimons johnsimons commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

What this fixes

Checking the audit instance's EF review fixes against the primary turned up four bugs in the primary instance's EF Core persisters. The EF Core persisters are not released yet, so no existing data needs migrating.

Time range filters return a 500 on PostgreSQL

FilterBySentTimeRange serves the from and to parameters of GET /api/messages2. FilterByLastModifiedRange serves the modified parameter of /api/errors, /api/endpoints/{name}/errors and the failure group routes. Both passed the parsed DateTime straight into the query. A value with an offset parses as Kind=Local, and a value with no zone parses as Kind=Unspecified. Npgsql only writes UTC values to timestamptz, so both kinds threw an ArgumentException such as Cannot write DateTime with Kind=Local to PostgreSQL type 'timestamp with time zone', and the request returned a 500. Only Z values worked.

SQL Server was not affected, because UtcDateTimeConverter also converts query parameters. ServicePulse always sends toISOString(), so only other API clients hit this.

The EF filters now convert both ends of a range to UTC. A value with no zone is treated as UTC, as SQL Server already does. The shared DateTimeRange is unchanged, because the RavenDB persister uses it too.

Search cannot find non-ASCII header values

MessageHeaders serialized with the default encoder, which escapes every non-ASCII character. HeadersJson stored Bestellpr\u00FCfung for Bestellprüfung, so the full text index held the escape sequence. A search for the word found nothing on either provider.

Headers are now written with JavaScriptEncoder.Create(UnicodeRanges.All). Letters outside ASCII are stored as they are, while + & < > ' and emoji stay escaped as before. This encoder was picked over UnsafeRelaxedJsonEscaping because it fixes search without anyone having to reason about HTML safety.

A value that does not fit its column fails the whole ingestion batch

A 600 character MessageId or ConversationId failed the batch it arrived in. PostgreSQL raised 22001: value too long for type character varying(450), and SQL Server raised "String or binary data would be truncated". IngestionPipeline.Fail fails every message in the batch. After three attempts, ErrorIngestionFaultPolicy stores each message as a failed import, and re-importing such a message fails the same way.

  • MessageId has no index, so the column no longer has a length limit on FailedMessages or on FailedErrorImports. When GET /messages/{id}/body looks a message up by MessageId, it compares the full value.
  • ConversationId is indexed and keeps its 450 limit. A value of 450 characters or fewer is stored unchanged, so GUIDs and ordinary custom ids are not affected. A longer value is stored as its first 385 characters, then #, then the SHA-256 of the whole value as 64 hex characters, without splitting a surrogate pair. Two different long ids never share a stored value, and applying the function to a stored value returns it unchanged.
  • The function runs in an EF value converter on ConversationId. EF applies a converter to its own writes and to every query value compared with the property. The conversation lookup therefore finds a message by its full id or by its stored form, and a future query cannot forget the conversion.
  • Ingestion writes through hand-written SQL, which skips converters, so EFRecoverabilityIngestionUnitOfWork also applies the function next to TruncateTypeName.
  • ServicePulse shows the stored form of a long id, and the headers tab still shows the full value. In a deployment that mixes RavenDB and EF instances, a lookup by the stored form finds nothing on the RavenDB side.
  • Endpoint names, queue addresses and MessageType keep their limits. Transports keep endpoint names and queue addresses far below 450 characters, and TruncateTypeName already shortens MessageType.

A body with too many distinct words fails ingestion on PostgreSQL

The full text index is a GIN index on a to_tsvector expression over the headers, the body and the message type. to_tsvector fails once the vector passes 1 MB. Because the index is on an expression, the error surfaces on INSERT and fails the batch. A 1.35 MB body of 150,000 distinct words failed with 54000: string is too long for tsvector (1800526 bytes, max 1048575 bytes). BodyText holds at most MaxBodySizeToStore bytes, 100 KB by default, so the failure needs that setting raised to around 1 MB.

The index now covers substring(COALESCE(body_text, ''), 1, 262144), and the search dialect renders exactly that from (message.BodyText ?? "").Substring(0, 262144). Words past the first 262,144 characters of a body are no longer searchable on PostgreSQL. SQL Server fills its full text index in the background, so a large body cannot fail an insert there, and its index is unchanged.

The index expression and the query expression must match character for character, or PostgreSQL falls back to a sequential scan without any error. FullTextSearchIndexTests pins the two together. A one-off EXPLAIN of the search query, with sequential scans disabled, showed a Bitmap Index Scan on ix_failed_messages_full_text.

…e handling in PostgreSQL persistence

`to_tsvector` fails when the vector it builds exceeds 1 MB, so a message body with enough distinct words would fail its entire ingestion batch. The full text index now covers only the first 256 KB of the body, and the search query mirrors that with `Substring`. A new migration (`WidenMessageIdAndCapFullTextBody`) replaces the uncapped index with the capped one and drops the 450-character limit on `message_id` so long IDs are stored in full.

Long conversation IDs are truncated to fit the index using a prefix-plus-SHA-256 scheme, keeping messages in different conversations distinguishable even when they share a prefix. The same migration runs on SQL Server as `WidenMessageId` (no full text index change needed there).

Non-ASCII characters in headers were being escaped before indexing, so searches for words like "Bestellprüfung" never matched. The JSON serializer now writes letters outside ASCII as-is. Timestamps without a timezone offset are now treated as UTC so Npgsql does not reject them.
Comment thread src/ServiceControl.Persistence.EFCore/Infrastructure/FailedMessageQueryFilters.cs Outdated
Comment thread src/ServiceControl.Persistence.EFCore/EntityConfigurations/ColumnLengths.cs Outdated
JsonSerializer.Deserialize(headersJson, HeadersJsonContext.Default.DictionaryStringString) ?? [];
JsonSerializer.Deserialize(headersJson, context.DictionaryStringString) ?? [];

// The default encoder escapes every non-ASCII character, and full text search then indexes the escape sequence instead of the word. This one writes letters outside ASCII as they are, while HTML sensitive characters and emoji stay escaped.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why do html sensitive characters matter in a json serialisation?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They don't, for this column. System.Text.Json escapes them by default in case the JSON ends up inside HTML. HeadersJson never does: MessageHeaders.Read parses it, and the API serializes the headers again with its own settings.

The escaping also hurt search, because the full text index reads the stored text. Create(UnicodeRanges.All) still wrote ' as \u0027, so "The given key 'CustomerId' was not present in the dictionary." was indexed with the token u0027customerid. A search for CustomerId found nothing on either provider. The same happened to a nested type name after a +.

Headers now use JavaScriptEncoder.UnsafeRelaxedJsonEscaping. A new test searches for a word in single quotes on both providers.

…uoted-word search

Move `UtcDateTimeConverter` to the shared `ServiceControlDbContext` so both SQL Server and PostgreSQL apply it to query parameters, removing the need to manually call `AsUtc` before filtering on date ranges. Timestamps without a zone are now treated as UTC rather than rejected by Npgsql.

Switch the over-length conversation ID separator from `#` to `~` so truncated IDs need no URL encoding when ServicePulse embeds them in a URL path.

Switch JSON header encoding to `UnsafeRelaxedJsonEscaping` so apostrophes, plus signs, and non-ASCII letters are written literally; the previous encoder caused full text search to index escape sequences instead of the words they represented.
@johnsimons
johnsimons merged commit eabac22 into master Sep 30, 2026
69 of 70 checks passed
@johnsimons
johnsimons deleted the john/ef_primary_fixes branch September 30, 2026 08:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants