Appier Studies How AI Detects Missing Answers and Selects Reasoning Languages
Appier detailed two studies in a press release on improving the reliability of AI agents. The research examined whether large language models can detect insufficient information and how their reasoning language affects performance.
In one study, researchers tested 28 models on multiple choice questions. Accuracy fell by 30% to 50% when "none of the above" was correct, as models often chose an incorrect option instead. Training with direct preference optimization improved their ability to identify questions with no valid answer by nearly 30 percentage points.
The second study found that models often reasoned in languages with more training resources, such as English, even when prompted in another language. For some models, the reasoning language differed from the response language in more than 90% of cases. Languages with more resources generally produced better results for mathematics and knowledge tasks, while reasoning in a local language performed better on cultural understanding and the detection of harmful or illegal requests.
Appier proposes routing tasks to a suitable reasoning language while returning answers in the user's preferred language. It also suggests that AI agents should check whether enough information is available before acting, then search again or ask a person when needed.
We hope you enjoyed this article.
Consider subscribing to one of our newsletters like AI Policy Brief or Daily AI Brief.
Also, consider following us on social media:
More from: AI Safety
Subscribe to AI Policy Brief
Weekly report on AI regulations, safety standards, government policies, and compliance requirements worldwide.
Whitepaper
Tensordyne Napier: What If One Rack Could Do the Work of Nine?
Tensordyne
This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.
Read more