Gathering search data, training a model, and reading pages requested by the user are different purposes. OpenAI's OAI-SearchBot, GPTBot, and ChatGPT-User also provide guidance on purpose and control methods.
In addition to robots.txt, check whether authentication, server responses, or security systems restrict access. Allowance is determined based on the site's content usage policy and the purpose of search participation.
There is no general rule that says you must allow training crawlers to be cited in searches. llms.txt is not an access-control mechanism or priority standard followed by every crawler.
