Who makes the tests? Who runs the tests? And who evaluates that the tests have meaning? As long as it is the AI, or you (with your self-admitted limited experience), how can you be sure it is meaningful?
Yeah, but you're missing a gut intuition if something is off.
I wrote a fancy polygon decomposition algorithm in university (pre-AI) which my professor didn't seem very impressed by because it was missing some sort of mathematical rigor. Yet everything I threw at it worked! Even he couldn't find a counter example.
It took a while for me to find some failing cases but it turned out they did exist.
But hey, maybe all I was missing is an AI-written lean proof.