Use upper-case HTS predicates to match the functional index - #758
Conversation
Replace lower() with upper() on both sides of handwritten JPQL comparisons in UserTableHtsJdbcRepository, preserving entity-type filtering and case-insensitive filter and LIKE behavior. The idx_user_table_upper_db_table index on (upper(database_id), upper(table_id)) cannot match lower() expressions. At roughly 6,000 QPS, production EXPLAIN ANALYZE measured 513,418 scanned rows / ~495ms before the change and an index lookup of 1 row / ~0.041ms with upper(). Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Thanks for the fix @ruolin59. Do we need index on the entity_type column? |
no, that would be the incorrect thing to do here. the issue isn't the lack of index on |
Replace the derived, unverified definitions of user_table_row and soft_deleted_user_table_row with their production SHOW CREATE TABLE output (minus the entity_type column, which is added by 0001). For user_table_row this closes several gaps that a derived definition could not capture: the idx_user_table_upper_db_table functional index on upper(database_id)/upper(table_id) that the service's query plan depends on, the InnoDB engine and utf8mb4 / utf8mb4_0900_ai_ci charset and collation, and the physical column order. It also fixes the column types (database_id, table_id to varchar(255), metadata_location to varchar(255)), makes version nullable, adds the table_version and deleted_ts columns, and drops last_modified_time, which production does not have. soft_deleted_user_table_row gains its InnoDB engine and utf8mb4 / utf8mb4_0900_ai_ci charset and collation; it has no secondary index. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
#758) (#761) ## Summary Reverts two related commits: - #696 "Add entityType discriminator and table-scoped HTS queries" - #758 "Use upper-case HTS predicates to match the functional index" #758 was a hotfix for a production incident (HTS connection saturation) caused by `getUserTable` scanning the full `user_table_row` table instead of using the `idx_user_table_upper_db_table` functional index, a regression introduced by #696's handwritten `lower(...)` predicates. This PR reverts both changes back to the pre-#696 state, restoring the original `findByDatabaseIdIgnoreCaseAndTableIdIgnoreCase`-based query path (which naturally matched the functional index) and removing the `entity_type` discriminator column, `EntityType` enum, and related JDBC/API/test surface added by #696. A follow-up PR will reintroduce both changes together, combined into a single commit, so the entityType feature and its required index-compatible predicate fix land atomically. ## Test plan - `./gradlew :services:housetables:test :services:common:test` passes. - `./gradlew spotlessCheck` passes (run with `-x CopyGitHooksTask`, a pre-existing worktree-incompatibility in the git-hooks Gradle task, unrelated to this change). Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…#696, linkedin#758 combined) Combines two previously-separate commits into one so the entityType discriminator feature and its required index-compatible predicate lands atomically: - Add entityType discriminator and table-scoped HTS queries (originally linkedin#696) - Use upper-case HTS predicates to match the functional index (originally linkedin#758, a hotfix for a production incident caused by linkedin#696's lower(...) predicates bypassing idx_user_table_upper_db_table) See the reverted PR (linkedin#761 revert of linkedin#696/linkedin#758) for background on why these two were split apart and are now being reintroduced together. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…fix (linkedin#696, linkedin#758)" (linkedin#761) This reverts the revert in linkedin#761, restoring linkedin#696 and linkedin#758 combined into a single commit: - Add entityType discriminator and table-scoped HTS queries (originally linkedin#696) - Use upper-case HTS predicates to match the functional index (originally linkedin#758, a hotfix for the production incident caused by linkedin#696's handwritten lower(...) predicates bypassing the idx_user_table_upper_db_table functional index) Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Production incident
HTS in
prod-ltx1saturated and began refusing TCP connections at ~11:35 PDT on 2026-09-18, breaking Airflow partition sensors withConnection refused ... :4768.user_table_rowhas ~513,000 rows and a functional index:A functional index matches only the exact expression, so handwritten
lower(...)cannot use it.getUserTableruns at roughly 6,000 QPS; a full scan per request exhausted MySQL connections and HTS worker threads, and pods stopped accepting connections.What regressed
Before eca3ca3 (#696),
getUserTablewent throughfindById(key), adefaultmethod delegating to the Spring-derivedfindByDatabaseIdIgnoreCaseAndTableIdIgnoreCase. Spring Data emitsupper(...)forIgnoreCase(JpaQueryCreator$PredicateBuilder.upperIfIgnoreCase), so it matched the index. That is why the index is onupper.#696 replaced that call with a new explicit
@Queryusing handwrittenlower(...), following the file's existing convention for hand-written queries. The method name still reads like the derived finder it displaced, which is how it passed review.The six pre-existing
lower(...)queries were all on cold paths — filters, search, rename. They had been scanning for a long time without anyone noticing, because a 500ms scan at low QPS is invisible. Moving the hot path onto that convention is what turned it into an outage.The fix
lower(→upper(at all 30 token occurrences across the 15 comparison sites inUserTableHtsJdbcRepository, both sides of every comparison. Nothing else.Production
EXPLAIN ANALYZE, measured on the real table:Table scan on user_table_rowIndex lookup using idx_user_table_upper_db_tableThe
entity_typepredicate is not the cause. In the fixed plan it is demoted to a cheap residual filter above the index lookup. It is implicated only because it arrived in the same commit.Why this is semantically safe
Production collation is confirmed
utf8mb4_0900_ai_ci, which is case- and accent-insensitive. Under it,lower(a) = lower(b),upper(a) = upper(b)anda = bare equivalent. The wrapping affects index eligibility, not which rows match.At the sole
LIKEsite both column and pattern fold with the same function, and%,_and escapes are non-alphabetic soupper()leaves them byte-identical — wildcard semantics are unchanged.Tests
Two tests added to
HtsRepositoryTest, covering cross-case matching for the filter andLIKEfamilies. Those families were previously exercised only with same-case data, so a botched substitution could have slipped through; point reads, rename and deletes already had cross-case coverage.These tests pass both before and after the change, and that is deliberate. There is no red phase because the change is behaviour-neutral by design.
What the tests prove: behaviour preservation across case for the affected query families.
What they cannot prove: index selection, scan avoidance, or latency. Tests run against H2 in MySQL mode, which has no functional indexes and no meaningful planner. The performance claim rests solely on the production
EXPLAIN ANALYZEabove.Suites pass on JDK 11: housetables 406, common 12, zero failures, errors or skips.
Deliberately out of scope
SoftDeletedUserTableHtsJdbcRepositoryhas 24lower(tokens across 12 lines onsoft_deleted_user_table_row. That is a different physical table whose indexes are unconfirmed — flipping it blind could be a no-op or a pessimisation. Its paths are cold (restore, purge, querying deleted tables). Needs its ownSHOW INDEXbefore anyone touches it.The schema record is corrected in this PR.
services/housetables/ddl/0000__baseline.sqlpreviously recorded onlyPRIMARY KEY (database_id, table_id)foruser_table_row, and its own header warned that a derived definition "cannot capture secondary indexes". So the index this fix depends on was documented nowhere, and anyone reconstructing the table from that file would have reintroduced this outage.user_table_rowandsoft_deleted_user_table_roware now transcribed from productionSHOW CREATE TABLE. Foruser_table_rowthat closed more than the index:database_id/table_idwere recorded asvarchar(128)but arevarchar(255),metadata_locationasvarchar(512)but isvarchar(255),versionasNOT NULLbut is nullable,last_modified_timewas recorded but does not exist, andtable_versionanddeleted_tsexist but were absent. Engine, charset and collation were missing from both tables.soft_deleted_user_table_rowhas no secondary index, and that is now recorded as the real state rather than an omission.job_rowandtable_toggle_ruleare deliberately untouched — no production output was available for them, and guessing would recreate exactly the failure this PR is fixing. The header now says which two tables are verified and which two are not.This also explains why no local or containerised MySQL could have caught the regression: the
oh-only-mysqlrecipe bootstraps from this DDL, so a local database had no functional index andlower()versusupper()was indistinguishable there.