Skip to content

MDEV-41303: rand() in a semi-join subquery is checked on outer rows - #5799

Open
DaveGosselin-MariaDB wants to merge 1 commit into
11.4from
11.4-mdev-41303-no-semijoin-for-rand
Open

DaveGosselin-MariaDB wants to merge 1 commit into
11.4from
11.4-mdev-41303-no-semijoin-for-rand

Conversation

@DaveGosselin-MariaDB

Copy link
Copy Markdown
Member

Converting an IN subquery to a semi-join moves its WHERE into the parent WHERE. A condition there such as rand(1) < 0.09 reads no table, so it is attached to the last table of the join order that is outside any materialized semi-join. With SJ-Materialization it was checked once for each outer row instead of once for each row of the subquery, and the query returned a wrong count.

Do not convert a subquery to a semi-join when it contains a function with a random result (UNCACHEABLE_RAND). Derived tables already follow this rule. ROWNUM sets the same flag, so this replaces the check for ROWNUM.

Converting an IN subquery to a semi-join moves its WHERE into the
parent WHERE.  A condition there such as rand(1) < 0.09 reads no table,
so it is attached to the last table of the join order that is outside
any materialized semi-join.  With SJ-Materialization it was checked once
for each outer row instead of once for each row of the subquery, and
the query returned a wrong count.

Do not convert a subquery to a semi-join when it contains a function
with a random result (UNCACHEABLE_RAND).  Derived tables already follow
this rule.  ROWNUM sets the same flag, so this replaces the check for
ROWNUM.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@spetrunia

Copy link
Copy Markdown
Member
explain extended
select count(*) from t1
where t1.a in (select c from t2 where rand(1) < 0.09);

Before the patch, we get:

Note	1003	select count(0) AS `count(*)` 
from `test`.`t1` semi join (`test`.`t2`) where rand(1) < 0.09

After the patch, we get (I've added formatting):

Note	1003	/* select#1 */ select count(0) AS `count(*)` 
from  
  <materialize> (/* select#2 */ select `test`.`t2`.`c` from `test`.`t2` where rand(1) < 0.09) 
  join `test`.`t1` 
where 
  `<subquery2>`.`c` = `test`.`t1`.`a`

So it is still converted into a semi-join.
But now the semi-join is a non-merged semi-join.
Subquery's WHERE condition stays in the subquery, so conversion to non-merged semi-join is fine

2 MATERIALIZED t2 ALL NULL NULL NULL NULL 80 100.00
Warnings:
Note 1003 select count(0) AS `count(*)` from `test`.`t1` semi join (`test`.`t2`) where rand(1) < 2
set optimizer_switch='firstmatch=default';

@spetrunia spetrunia Oct 1, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does re-running the test with firstmatch enabled add any value?
The first query uses the same query plan.
The second uses First Match, but we've already saw

from `test`.`t1` semi join (`test`.`t2`) where rand(1) < 2

above...

set optimizer_switch='firstmatch=off';

# The subquery must check rand(1) < 0.09 on the 80 rows of t2. Checking
# it on the 4 rows of t1 instead rejects every row.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please use --echo # so that this shows up in .result too.
the last line should be something that is clear about the expected output, like

-echo # Must produce count(*)=1,  not 0. 
# it on the 4 rows of t1 instead rejects every row.
select count(*) from t1
where t1.a in (select c from t2 where rand(1) < 0.09);
explain extended

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same here

--echo # Warning must have "<materialize>(...) JOIN..."  and not "`test`.`t1` semi join (`test`.`t2`)" 
@spetrunia

spetrunia commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Remember I've asked this question on the call:

If the SELECT has an Item with 'item->used_tables() & RAND_TABLE_BIT, will that translate into select_lex->uncacheable & UNCACHEABLE_RAND` for the SELECT that that Item is located

Claude gives this answer:

Items that have RAND_TABLE_BIT but never set UNCACHEABLE_RAND:

  • Item_func_get_user_var: used_tables() returns the bit when the item isn't const (item_func.h:3672).
  • Item_func_sysdate_local: item_timefunc.h:961.
  • Item_func_sp: the bit is set when the stored routine is non-deterministic (item_func.cc:6951).
  • Item_insert_value (item.h:7460).
  • A not-fully-parsed Item_default_value (item.cc:10207).
  • Item_func_xml_* (item_xmlfunc.cc:197).
  • Item_direct_view_ref in HAVING, (hallucination removed), the Item_sum / window-function caches. These add the bit only as a "don't push or move" marker.

I've verified it for SYSDATE()...
I don't see any items in the above list that would allow to construct a failing testcase, though...
But this might be something to keep in mind for MDEV-41146.

@DaveGosselin-MariaDB

Copy link
Copy Markdown
Member Author

Remember I've asked this question on the call:

If the SELECT has an Item with 'item->used_tables() & RAND_TABLE_BIT, will that translate into select_lex->uncacheable & UNCACHEABLE_RAND` for the SELECT that that Item is located

Yes, I investigated with Claude and discovered that the answer to this is 'no'. I posted this PR because the answers to the three questions we discussed on the call indicated that this solution is sound, so I thought I would post it. I planned to discuss the questions during the team call today.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

2 participants