AI Assistants for Beginners: What They Can and Can't Do for You
AI assistants for beginners: in a 758-consultant experiment, AI help improved work on 18 tasks but cut correct answers on another. What they do well and badly.

An AI assistant can make your work better on one task and worse on the next, and its answers look equally convincing both times. In a 2023 experiment with 758 Boston Consulting Group consultants, those randomly given GPT-4 did better-rated work on 18 tasks the tool handled well, but on a task beyond its abilities, a separate group of AI users was 19 percentage points less likely than consultants without it to reach the right answer.1
The first rule of AI assistants for beginners follows from that result: let them draft, rewrite, explain and brainstorm, and check anything that states a fact. They produce answers by predicting likely text, so they can invent details and sources, miss recent events and agree with you too readily. Judging the answer stays your job.
What an AI assistant is, and where its answers come from
An AI assistant is a chatbot built on a large language model. A 2024 framework from the US National Institute of Standards and Technology (NIST) explains that such models generate text by predicting the next word from statistical patterns in their training data, which can produce accurate, consistent answers or wrong and internally inconsistent ones.2
NIST calls confidently stated false content confabulation; most people call it hallucination.2 Language models produce plausible but incorrect content because they return the most likely output for an input without actually understanding it, notes the UK government’s 2025 AI Playbook, which also warns that they are not a substitute for professional advice in legal, medical or other critical areas.3
What an assistant knows can also be out of date. By 2025 some well-known assistants could pull live results from the web, but many language models had no real-time internet access and knew only what was in their training data, the Playbook adds.3 Before relying on an answer about anything recent, check whether the assistant searched, and open what it found.
AI assistants for beginners: the jobs they do well
AI assistants do best at drafting, suggesting replies and generating ideas: a 2023 writing experiment in Science, a 2025 study of support agents in the Quarterly Journal of Economics and a 2024 story-writing experiment in Science Advances all measured gains on those jobs.456
In the writing experiment, which was preregistered and run online, 453 college-educated professionals did writing tasks from their own fields, and those randomly given ChatGPT took 40 percent less time, with quality rated 18 percent higher.4 In the authors’ earlier working paper on the same experiment (not peer reviewed), the tasks, such as press releases, short reports and delicate emails, took 20 to 30 minutes and needed little company-specific knowledge, which the authors noted may inflate the benefit.7
The least experienced workers often gain the most. In the support-agent study, which covered 5,172 agents at one software company, a tool that suggested replies, rolled out in 2020 and 2021, raised issues resolved per hour by 15 percent on average. Less experienced and lower-skilled agents improved in both speed and quality, while the most experienced saw small gains in speed and small declines in quality. The tool was rolled out in stages, with only a small randomized pilot, and the researchers worked with the vendor that supplied it.5 The writing experiment likewise narrowed the gap between stronger and weaker performers.4
Assistants can also feed you ideas, at some cost to variety. In the story-writing experiment, access to story ideas from GPT-4 made short stories more creative, better written and more enjoyable in evaluators’ eyes, especially for less creative writers, but the AI-assisted stories were more similar to one another than stories written without it.6
Better work on 18 tasks, worse answers on one
A randomized experiment run with Boston Consulting Group and published in Organization Science in 2026 measured both what an AI assistant can and can’t do, using the same tool at the same firm. It improved one group of consultants’ work on 18 tasks and made a second group less accurate on a task designed to be of comparable complexity.1
The study
Moderate evidence
Same tool, opposite results for 758 consultants (Organization Science, 2026)
Researchers from Harvard Business School, Wharton, MIT and Warwick, working with Boston Consulting Group, randomly gave consultants no AI, GPT-4, or GPT-4 plus an overview of prompting. On 18 product-development tasks, AI users completed more tasks, and faster, and graders scored their work 30 to 34 percent higher. A second group tackled a business case where the spreadsheet looked complete but the interview notes held the details that mattered: 84.5 percent of those without AI reached the right recommendation, against 70.6 percent with GPT-4 alone and 60 percent with GPT-4 and the overview.1
Wrong answers could still look convincing: graders who did not know the correct solution rated the AI users’ recommendations as better argued, whether or not they were right. The main caveats are one firm, consultants early in their careers, a 2023 model and a single task outside the AI’s range, which the authors call an inherent limitation.1
The authors call this uneven boundary a jagged technological frontier: tasks that seem equally hard to a person can fall on opposite sides of it, and a worker may not know in advance which side a given task is on.1
- Inside the frontier: tasks where AI help made consultants’ work better and faster
- Outside it: a task of similar difficulty where consultants using AI were more often wrong, though their answers still read as convincing
A broader review points the same way. A 2024 meta-analysis of 106 experimental studies in Nature Human Behaviour found that people working with AI beat people alone, but on average did worse than whichever of the two, person or AI, was better on its own; the combinations lost ground on decision tasks and did relatively better on tasks that involved creating content. About 85 percent of the results came from decision tasks, such as choosing among set options, in studies published from 2020 to mid-2023.8
Where assistants go wrong: confident errors, weak sourcing and flattery
In a 2025 study coordinated by the European Broadcasting Union and the BBC, journalists at public broadcasters in 18 countries found at least one significant issue in 45 percent of the answers that four free AI assistants gave to news questions. The problems ranged from outdated or inaccurate details to sources that did not back the claim.9
45%of 2,709 answers to news questions from four free AI assistants had at least one significant issue, in a 2025 audit by 22 public broadcastersSource: EBU and BBC, 2025Sourcing was the biggest problem, affecting 31 percent of answers; 20 percent had significant accuracy problems, with outdated information among the most common. The assistants almost never declined: only 17 of 3,113 questions were refused, and the participating broadcasters said the confident tone and the inability to express uncertainty made errors worse.9
The four assistants differed widely, with significant issues in 30 to 76 percent of their answers, so it pays to try more than one on a task whose answer you already know. The broadcasters have a stake in how AI handles their journalism, and the EBU and BBC published the report themselves rather than in a peer-reviewed journal.9
Summaries carry a quieter risk. When 10 widely used models summarized abstracts and articles from leading science and medical journals, in a 2025 study in Royal Society Open Science, most stated the findings more broadly than the originals did, often by leaving out details that limited the conclusions.10
- Myth
- An answer that sounds sure of itself and cites a source is probably right.
- Fact
- Confidence is not evidence. In a 2025 audit by 22 public broadcasters, four AI assistants refused only 17 of 3,113 news questions, and 31% of answers had significant sourcing problems.
Assistants also lean toward the person asking, a tendency researchers call sycophancy. In a peer-reviewed 2024 study by researchers at the AI company Anthropic, five assistants tested in 2023 gave more positive feedback on arguments the user said they liked, and often abandoned correct answers when the user pushed back.11
Further reading
Co-Intelligence: Living and Working with AI
Ethan Mollick, a co-author of the consultant experiment in this article, on the jagged frontier and how to work alongside AI assistants.
As an Amazon Associate WiserHours earns from qualifying purchases.
Six habits for using an AI assistant without being misled
Six habits follow from randomized experiments, audits of AI answers and UK and Australian government guidance: start with work you can check, give the assistant your material, verify what it states, keep your own view out of requests for feedback, keep sensitive data out of the chat, and time the results. They are ordered by the strength of the evidence behind them, strongest first.
1. Start with work you can check yourself
In the consultant experiment, AI help raised quality on tasks inside the tool’s range, while outside it the AI’s output was inaccurate and dragged human performance down.1 Begin with drafts, rewrites, outlines and explanations of material you know well, where a wrong turn is easy to spot, and move to harder work as you learn where your assistant slips.
The matrix is our editorial rule of thumb rather than a tested method: it combines the consultant findings with the UK Playbook’s warning about legal, medical and other critical areas.13
2. Give it the source material, then compare
Paste in the document or data you are asking about, so you can check each claim in the answer against the original. Compare most carefully who and what a finding applies to: in the 2025 summary study, models tended to drop exactly the details that limit a conclusion.10
3. Check facts, figures, quotes and links before you use them
Open each cited link and confirm it says what the answer claims, and confirm names, numbers, dates and quotations at the original source. In the broadcasters’ audit, 12 percent of the answers that included a direct quote had significant problems with it, and some quotes appeared to be fabricated.9
4. Keep your opinion out of requests for feedback
Ask what is weak about a draft before saying you like it. In the sycophancy study, assistants gave more positive feedback on arguments the user said they liked and more negative feedback on arguments the user disliked, and they often changed correct answers when challenged.11 If an assistant reverses itself after you push back, treat the change as a reason to check, not as a correction.
5. Keep sensitive and personal information out
The UK’s National Cyber Security Centre advised in 2023 against including sensitive information in queries to public chatbots, noting that providers store queries, may be able to read them and are likely to use them in developing their services.12 Australia’s privacy regulator recommends that organizations not enter personal information, particularly sensitive information, into publicly available AI tools.13 A 2025 Stanford analysis of six large US developers’ privacy policies, as they stood in May 2025, found that all six appeared to train their models on users’ chats by default, and most required users to opt out.14 Check your assistant’s data settings, and at work follow your employer’s rules on using AI at work.
6. Time a task with and without it
In a 2025 randomized trial by the research nonprofit METR, 16 experienced open-source developers expected AI tools to cut their task time by 24 percent and afterwards believed they had saved 20 percent, but tasks where AI was allowed took 19 percent longer.15 METR said in 2026 that developers were probably being sped up more by then, while warning that developers’ own estimates of their speedup can be unreliable.16 Larger trials run by employers point the other way: across three field experiments at Microsoft, Accenture and another large company, starting in 2022 and reported in Management Science in 2026 by a team including two Microsoft researchers, thousands of developers using an AI coding assistant completed about 26 percent more tasks, an imprecise estimate, with bigger gains for less experienced developers.17 Do a routine task both ways a few times and compare the clock and the result, not your impression.
Checks for any AI answer you plan to use
How reliable is the research on AI assistants?
The firmest findings on AI assistants come from randomized experiments: they speed up professional writing tasks, and their help is uneven from task to task. The error rates come from audits and controlled tests of the models, and the privacy advice from government guidance. Company-run research is marked as such, and every study here tested AI systems from 2025 or earlier.
| What it is | What the best evidence found | Evidence |
|---|---|---|
| Professional writing tasks | 40% less time, quality rated 18% higher; 453 professionals, occupation-specific writing tasks | Trial, moderate4 |
| Reply suggestions for support agents | 15% more issues resolved per hour; less experienced agents gained in speed and quality | Observational (staged rollout), moderate5 |
| Story ideas | Stories rated more creative, especially for less creative writers, but more alike | Trial, limited (one study, 293 writers)6 |
| Consulting tasks inside and outside the AI’s range | Quality up 30 to 34% on 18 tasks; fewer right answers on 1 task beyond the AI’s range | Trial, moderate1 |
| People working with AI, across 106 studies | Better than people alone, worse than the better of person or AI alone; losses on decision tasks | Meta-analysis, moderate8 |
| Answers to news questions | 45% had a significant issue; 31% had sourcing problems | Observational (audit), moderate9 |
| Summaries of research | Most of 10 models stated findings more broadly than the originals | Observational (model test), moderate10 |
| Agreeing with the user | Five assistants shifted feedback and answers toward the user’s view | Observational (model test), moderate; company research11 |
| Feeling faster | 16 developers believed AI saved them 20% of their time; tasks took 19% longer | Trial, limited (preprint)15 |
The bottom line
An AI assistant is a fast, fluent first-draft partner that cannot tell you when a task has crossed from its strengths into its weaknesses.1 Start where you can check the work, give it your material, verify every fact, figure and source before it leaves your desk, and keep anything sensitive out of the chat.
Frequently asked questions
Does an AI assistant understand what it writes?
No, not in the way a person does, according to the UK government's 2025 AI Playbook. Language models return the most likely output for a given input, based on patterns in their training data, without actually understanding the content, and although some appear to reason, they are in no way sentient. That is why a fluent, well-organized answer can still be wrong.
Does prompt training make an AI assistant's answers more reliable?
Not in the one large experiment that tested it. In the Boston Consulting Group experiment published in 2026, consultants who also received an overview of prompting techniques gained somewhat more on tasks the AI handled well, yet on the task beyond its abilities, those given the overview were, if anything, more often wrong than those given the tool alone. The authors report that difference as significant only at the 10 percent level.
Are AI assistants getting more accurate?
Somewhat, but unevenly. In the BBC's checks of assistant answers about its own journalism, the share with significant issues fell from 51 percent in late 2024 to 37 percent in mid-2025, yet the wider 2025 study by 22 public broadcasters still found errors across every assistant and language. A 2025 study of research summaries by 10 models found that newer models overgeneralized findings more often than older ones.
Does telling an AI assistant to be accurate make it more accurate?
Not reliably. In a 2025 Royal Society Open Science study of 10 language models summarizing scientific abstracts and articles, a prompt that explicitly asked the models to avoid inaccuracies made overgeneralized summaries more likely than a simple request to summarize. Instructions do not replace comparing an answer with its source, especially the details about who and what a finding applies to.
Sources
- Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality. Dell'Acqua, F., McFowland III, E., Mollick, E., et al. (2026). Organization Science, 37(2)
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1). National Institute of Standards and Technology, US Department of Commerce (July 2024)
- Artificial Intelligence Playbook for the UK Government. Government Digital Service, UK (10 February 2025)
- Experimental evidence on the productivity effects of generative artificial intelligence. Noy, S. & Zhang, W. (2023). Science, 381(6654)
- Generative AI at Work. Brynjolfsson, E., Li, D. & Raymond, L. R. (2025). The Quarterly Journal of Economics, 140(2)
- Generative AI enhances individual creativity but reduces the collective diversity of novel content. Doshi, A. R. & Hauser, O. P. (2024). Science Advances, 10(28)
- Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence (working paper, not peer reviewed). Noy, S. & Zhang, W. (2 March 2023). Working paper, MIT Department of Economics
- When combinations of humans and AI are useful: A systematic review and meta-analysis. Vaccaro, M., Almaatouq, A. & Malone, T. (2024). Nature Human Behaviour, 8(12)
- News Integrity in AI Assistants: An international PSM study. European Broadcasting Union and BBC (October 2025)
- Generalization bias in large language model summarization of scientific research. Peters, U. & Chin-Yee, B. (2025). Royal Society Open Science, 12(4)
- Towards Understanding Sycophancy in Language Models. Sharma, M., Tong, M., Korbak, T., et al. (2024). International Conference on Learning Representations (ICLR 2024)
- ChatGPT and large language models: what's the risk? National Cyber Security Centre, UK (14 March 2023)
- Guidance on privacy and the use of commercially available AI products. Office of the Australian Information Commissioner (published 21 October 2024, updated 17 January 2025)
- User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies. King, J., Klyman, K., Capstick, E., Saade, T. & Hsieh, V. (2025). Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 8(2)
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. Becker, J., Rush, N., Barnes, E. & Rein, D. (2025). METR, arXiv preprint 2507.09089
- We are Changing our Developer Productivity Experiment Design. Becker, J., Rush, N., Cunningham, T., Rein, D. & Mahamud, K. (2026). METR, 24 February 2026
- The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. Cui, K. Z., Demirer, M., Jaffe, S., Musolff, L., Peng, S. & Salz, T. (2026). Management Science, articles in advance
How we researched this
We searched the web, Crossref, PubMed and arXiv in September 2026 and read guidance from NIST, the UK government, the UK National Cyber Security Centre and Australia's privacy regulator. We preferred peer-reviewed randomized experiments and meta-analyses, and we label company-run research, industry research, preprints and working papers as such. Sources date from 2023 to 2026. Main limitation: most studies tested AI models from 2023 to 2025, and results can change as models do.




