What Large Language Models Can and Cannot Do for Marketing

Marketer fact-checking language model drafts

Large language models have become a practical part of marketing work faster than many previous software categories. They are now built into search platforms, office suites, creative tools, customer service products, analytics interfaces, and marketing technology stacks. For many teams, the question is no longer whether they have encountered these systems. It is whether they understand what they are actually good at, where they fail, and what kind of oversight marketing work requires when language generation becomes easy and cheap.

That distinction matters because large language models, or LLMs, are often discussed as if they were general-purpose thinking systems. In practice, they are better understood as statistical language systems trained to predict the next token, usually a word fragment, in a sequence based on patterns learned from very large datasets. That sounds technical, but the practical implication is straightforward. These models are designed to generate plausible language. They are not designed to guarantee truth, maintain a verified understanding of the world, or apply judgment the way an experienced strategist, copy editor, attorney, researcher, or brand leader would.

For marketing professionals, that is both the source of their value and the source of their risk.

What large language models actually do

An LLM is trained on large collections of text and learns relationships among words, phrases, structures, and contexts. During use, it receives a prompt and generates output one token at a time by estimating which next token is most likely to fit the context. Through additional tuning methods, including instruction tuning and reinforcement learning from human feedback, many consumer-facing systems have become better at following requests, summarizing information, changing tone, extracting patterns from text, and carrying on a useful dialogue.

This architecture makes LLMs particularly good at tasks where the underlying problem is linguistic patterning. They can draft headlines, reorganize notes, turn a long transcript into a concise brief, propose alternative subject lines, summarize customer reviews, classify themes in open-ended responses, convert rough ideas into clearer prose, and adapt copy to multiple formats.

These are not minor capabilities. A large share of marketing work is language work, including internal language. Teams write campaign briefs, competitive summaries, media rationales, audience descriptions, FAQ pages, pitch materials, social copy, ad variants, presentation outlines, and research summaries. In many of those settings, speed and iteration matter. LLMs can compress the time needed to move from blank page to workable draft.

That is why they are proving useful in real workflows. A model does not need to be a reliable expert in order to be useful as a drafting and synthesis system. It only needs to reduce low-value effort without introducing unacceptable error.

Why they are strong at drafting and synthesis

The most reliable marketing use cases tend to align with what the models are structurally good at.

First, LLMs are effective at rewriting and reformatting. If a team already has source material, such as product documentation, webinar transcripts, customer service logs, press releases, creative briefs, or research notes, a model can often convert that material into different forms quickly. It can shorten, expand, simplify, localize, and reorganize text with reasonable fluency.

Second, they are useful for synthesis. Given a bounded set of material, an LLM can often identify recurring themes, extract high-level takeaways, summarize objections, cluster customer pain points, or create a first-pass outline of a report. This can help researchers, strategists, and account teams process large volumes of language faster than manual review alone.

Third, they are useful for variation at scale. Marketing often requires multiple executions that express the same message differently across channels, audience segments, and creative tests. LLMs can generate many alternatives more quickly than a human team starting from scratch each time. That can help with ideation and operational throughput, especially when outputs are reviewed and narrowed by people who understand the brand and the audience.

None of this means the model understands the campaign, the customer, or the business objective in a human sense. It means the model is good at producing linguistically coherent responses that resemble the kinds of text found in its training and fine-tuning data.

That distinction is central to using the technology responsibly.

Where the problems begin: hallucination and unsupported claims

The most widely discussed limitation of LLMs is hallucination, the generation of content that is false, fabricated, or unsupported but presented in a confident and fluent way. This is not a rare edge case. It is a known property of current generative language systems.

Because the model is predicting plausible sequences rather than consulting a guaranteed factual database, it may invent statistics, misstate product features, fabricate source citations, attribute quotes incorrectly, or blend multiple concepts into a misleading answer. The output can sound polished enough to escape quick detection, which is what makes it risky in marketing contexts.

A fabricated historical claim in an internal brainstorming document may waste time. A fabricated product claim in a landing page, investor-facing deck, healthcare ad, or regulated-industry email campaign may create legal and reputational exposure.

The issue is not simply that a model can be wrong. Search engines, analysts, and junior staff can also be wrong. The issue is that the model produces error in the same fluent format as accurate material, often without clearly signaling uncertainty. That makes weak oversight particularly dangerous.

This is one reason leading AI providers and researchers have consistently documented hallucination as an unresolved limitation rather than a solved problem. Some approaches reduce the risk, including better prompting, retrieval systems, domain tuning, and tool use, but none eliminates it across all tasks.

Why stale knowledge still matters

Another practical limitation is that an LLM’s internal knowledge may be incomplete, outdated, or detached from recent developments. A base model is trained on data collected up to a certain point. If it is not connected to current sources during inference, it will not know about later events, policy changes, product updates, market shifts, campaign launches, or emerging cultural context unless that information is supplied in the prompt.

For marketers, this is not a technical footnote. It affects competitive analysis, trend interpretation, regulatory awareness, current-events sensitivity, product messaging, and channel planning.

A model may produce a tidy summary of a market category while missing major entrants or recent mergers. It may describe a social platform’s ad capabilities in outdated terms. It may overlook current privacy changes, creator economy shifts, or measurement constraints. If teams mistake fluent prose for current expertise, stale knowledge can quietly degrade decision quality.

This is especially relevant when LLMs are used for strategy documents, research summaries, or public-facing content that appears to rely on recent facts. Without source verification, it can be difficult to tell whether an answer reflects updated information, partial information, or no real grounding at all.

Source grounding is improving, but it is not automatic

One way organizations try to make LLM outputs more reliable is through retrieval-augmented generation, often called RAG. In a RAG system, the model is connected to a defined body of documents and can use retrieved material as context for generating an answer. In marketing settings, that document set might include brand guidelines, approved claims, knowledge bases, prior campaign reports, retailer requirements, legal language, product specs, or research summaries.

This can materially improve usefulness. If a model is drafting copy from approved source material rather than relying only on its training data, the risk of unsupported invention can decrease. It can also make outputs more organization-specific.

But source grounding should not be treated as a guarantee of truth. A retrieval system can pull the wrong document, retrieve incomplete context, miss relevant nuance, or give the model conflicting material. The model may still overstate what the source says, combine sources carelessly, or produce wording that sounds stronger than the evidence supports.

For marketers, the practical lesson is clear. “Connected to your data” does not mean “factually safe.” Grounded systems are generally more useful than ungrounded ones for enterprise work, but they still require editorial and subject-matter review.

Reasoning is uneven, not absent and not dependable

LLMs can appear to reason because they can produce step-by-step explanations, compare options, follow conditional instructions, and solve some structured problems. In many cases, especially routine ones, they are genuinely helpful for organizing decisions. They can sort messaging themes, identify gaps in a creative brief, propose test variables, or map customer objections to content ideas.

At the same time, their reasoning is inconsistent. A model may handle a complex prompt correctly in one instance and fail on a similar prompt in another. It may produce a persuasive explanation for a flawed conclusion. It may miss contradictions inside a brief, overlook quantitative errors, or follow superficial cues instead of business logic.

Researchers and model providers continue to improve performance on reasoning benchmarks, but benchmark gains should not be confused with dependable judgment in production marketing environments. A model that can generate a coherent media rationale or brand-positioning framework is not therefore qualified to decide whether the rationale is strategically sound, compliant, or aligned with market realities.

Marketing work often depends on constraints that are hard to infer from text alone: budget, channel economics, legal risk, retail relationships, internal politics, brand history, audience fatigue, category norms, and timing. LLMs can help structure those discussions, but they do not reliably resolve them.

What this means for common marketing use cases

The most effective use of LLMs in marketing usually comes from matching the tool to the task rather than asking it to do everything.

For content drafting, the models can be efficient starting-point tools. They are well suited to generating first drafts of blog outlines, social variants, email subject lines, product-description options, FAQ drafts, metadata, and campaign summaries. They can also help convert long-form material into shorter channel-specific assets. The oversight requirement here is brand, factual, and legal review. Fast drafting does not reduce the need for approval discipline.

For research support, LLMs can help synthesize survey verbatims, summarize interviews, cluster themes in support logs, and extract recurring questions from reviews or sales calls. This is often one of the more practical enterprise applications. The caution is that the model’s summary may flatten nuance, overstate consensus, or miss low-frequency but important signals. Researchers still need to inspect source material and validate interpretations.

For strategy and planning, LLMs can accelerate early-stage thinking by generating hypotheses, organizing brainstorms, identifying message territories, or turning raw notes into presentation-ready structure. Here the value is often administrative and combinatorial rather than decisional. A strategist may save time shaping a document, but the model is not a substitute for market knowledge or professional judgment.

For customer service and conversational interfaces, LLMs can improve response drafting and make chat interactions more flexible than older scripted bots. Yet this is also a high-risk area because fabricated policy statements, inaccurate troubleshooting, or mishandled edge cases can directly affect customer trust. Many organizations therefore use guardrails, limited domains, human escalation, and approved-response frameworks rather than fully open-ended generation.

For personalization, LLMs may help tailor messaging formats, summarize prior interactions, or generate more relevant content options. But personalization still depends heavily on data quality, segmentation logic, permissions, and channel execution. The language model is only one part of the system. It does not solve the underlying challenges of identity, consent, timing, measurement, or offer relevance.

What large language models do not change

The growing usefulness of LLMs does not eliminate several longstanding realities of marketing practice.

They do not replace the need for a clear strategy. A model can create many messages, but it cannot establish a sound market position on its own. If the segmentation is weak, the value proposition is unclear, or the brand voice is inconsistent, generated output will simply scale those problems.

They do not remove accountability for claims. Whether copy is drafted by a model, an agency, or an in-house team, the organization publishing it remains responsible for accuracy, substantiation, disclosures, and compliance.

They do not make source quality irrelevant. In fact, they often make it more important. Weak source material can be turned into polished but misleading output faster than before.

They do not remove the need for editing. Generated language may be grammatical and on-topic while still being generic, repetitive, tonally off, or strategically shallow. Many teams discover that the challenge shifts from writing every sentence manually to reviewing, correcting, and elevating machine-generated drafts.

They do not erase the importance of domain expertise. In regulated categories, technical industries, politically sensitive contexts, healthcare, finance, and brand-critical communications, expert review becomes more important, not less.

The operational question is not just output quality

Many early discussions about LLMs focused on whether the writing is good. For organizations, the more important question is whether the workflow around the writing is sound.

That includes governance over what tools are approved, what data can be entered into external systems, how prompts and outputs are stored, who reviews public-facing content, how model use is documented, and which tasks require human signoff. It also includes training staff to understand failure modes rather than simply teaching prompt tricks.

Several risks are operational rather than purely linguistic. Confidential campaign information may be pasted into systems without clear contractual protections. Teams may lose track of which outputs were source-checked. Junior staff may over-trust polished drafts. Managers may assume productivity gains are automatic while ignoring the time required for review and correction. Organizations may also underestimate how quickly inconsistency grows when different teams use different tools without shared standards.

These are management issues as much as technology issues.

Evidence of value is real, but it is uneven

There is credible evidence that generative AI tools can improve productivity in some knowledge-work tasks, especially drafting, summarization, and support for less experienced workers. But productivity gains vary widely by task design, user skill, domain complexity, and quality requirements. A tool that saves time on internal ideation may create more downstream work if outputs require heavy correction before publication.

Marketing leaders should therefore be cautious about generalized efficiency claims. A measured evaluation asks narrower questions. Does the model reduce time to first draft for routine campaign assets? Does it help research teams process open-ended responses faster without distorting findings? Does it improve customer service response speed while maintaining policy accuracy? Does it reduce production bottlenecks in localization or metadata generation?

Those are testable workflow questions. They are more useful than asking whether “AI can do marketing.”

How marketers should think about oversight

Oversight is not a ceremonial final glance at generated copy. It should be matched to the risk of the task.

Low-risk internal uses, such as meeting summaries, headline ideation, or draft outlines, may require only light review. Moderate-risk uses, such as blog drafts, email copy, and research summaries, generally require factual checking, brand review, and editorial revision. High-risk uses, such as product claims, regulated-industry messaging, executive communications, investor-facing materials, legal language, or crisis response, require much tighter controls and often should rely on preapproved sources or direct human authorship.

The review process should focus on the specific weaknesses of LLMs:

  • Verify factual claims, numbers, dates, quotations, and citations.
  • Check whether the output relies on current information or outdated assumptions.
  • Inspect whether conclusions are actually supported by the provided source material.
  • Review for subtle misreadings of brand tone, audience sensitivity, and legal constraints.
  • Look for generic language that sounds acceptable but says little.

This is less glamorous than broad claims about automation, but it is what determines whether the technology is useful in practice.

Why the distinction between language fluency and business judgment matters

The most important professional insight for marketers may be that language fluency is not the same as judgment. LLMs can speak in the form of expertise without possessing the accountability structures, evidence standards, contextual memory, or strategic responsibility that professional marketing work requires.

That does not make the tools trivial. It makes them specific.

They are strong assistants for drafting, restructuring, summarizing, and generating alternatives. They can reduce friction in content operations and help teams work through large amounts of text. They can support research synthesis and routine production. They can improve speed in places where the main problem is getting from material to draft.

They are not reliable standalone authorities for facts, sources, brand decisions, or strategic conclusions. Hallucination, stale knowledge, weak source grounding, and inconsistent reasoning are not side issues. They are defining operating conditions of the technology.

For advertising and marketing professionals, the practical path is neither blanket adoption nor blanket rejection. It is a clearer understanding of what these systems are actually doing when they produce language, where that is useful, and where professional oversight remains non-negotiable. Teams that understand that distinction will be in a better position to evaluate claims, design safer workflows, and use large language models where they genuinely improve the work rather than merely accelerating the appearance of it.

Leave a Reply

Discover more from American Advertising and Marketing Association | AAMA

Subscribe now to keep reading and get access to the full archive.

Continue reading