Skip to main content

AI-Generated Content and Copyright in 2026: What Small Businesses Can Actually Own, Use, and Risk

Published 15 min readMike ThriftMike Thrift
AI-Generated Content and Copyright in 2026: What Small Businesses Can Actually Own, Use, and Risk

You prompted an AI to write your product descriptions, design your logo options, and draft next week's email campaign — all before lunch. It felt like you finally have a marketing department without the headcount. Then a harder question lands: do you actually own what it made? Could a competitor copy it tomorrow with no recourse? Could you be infringing someone else's work without ever pasting a sentence?

In 2026 those questions stopped being hypothetical. A $1.5 billion settlement, a Supreme Court decision that closed one door on AI ownership, and a wave of federal rulings about how models are trained have rewritten the risk ledger for every business that uses generative tools. The good news is the pattern is now clear enough to act on, even while appeals are still pending.

This guide translates the 2026 legal landscape into what matters for your day-to-day decisions: what you can protect, where you are exposed, and the simple habits that keep your AI-assisted work both useful and defensible.

U.S. copyright protects "original works of authorship fixed in any tangible medium of expression." That phrasing hasn't changed, but its application to AI has finally been tested.

The U.S. Copyright Office launched a formal AI initiative in 2023 and has since released two reports focused on AI's impact. The second report put it plainly: copyright protection in the United States requires human authorship. The constitutional grant to secure rights for "authors" and the Copyright Act as interpreted by courts both point the same way. No court has recognized copyright in material created solely by a non-human system.

That principle was reinforced in early 2026 when the Supreme Court declined to review a case in which a researcher had listed a generative system as the sole author of an AI-generated artwork and sought registration. Lower courts had held that human authorship is a bedrock requirement and that listing a machine as author leaves nothing to transfer under work-for-hire. By denying certiorari on March 2, 2026, the Court left that requirement intact.

What this does not mean is that any use of AI poisons copyright. Works created with AI assistance can be protected where a human contributes meaningful creative choices — selecting, arranging, editing, combining, or transforming AI output in an original way. A prompt alone, even a detailed one, has consistently been viewed as an instruction to the system, not authorship of the output. The authorship is what you add after the model responds.

Takeaway for you: If you want to claim ownership, plan to show your human contribution. The prompt is the brief, not the work.

Training vs. Output: Two Different Fair Use Questions

Most 2026 rulings drew a bright line that matters directly to small businesses. How a model was trained, and what it outputs for you, are separate legal questions with different risk profiles.

Why training was called "spectacularly transformative"

In two major Northern District of California decisions in late June 2025, courts found that training a large language model on copyrighted books was transformative — one court called it "spectacularly so" and among the most transformative uses many judges will see in a lifetime. The reasoning: training ingests billions of tokens and adjusts internal weights to learn patterns of language, much like a person reading to learn how language works, without ever showing the training texts to users.

The same decisions emphasized that training copies were never displayed to end users and did not substitute for the market for the original books. A key factor in fair use — market effect — was found to favor the model developers where outputs did not replicate significant portions of the books and where the developers had added safeguards against regurgitation. Loss of a hypothetical licensing fee for training was not treated as cognizable harm where the use was deemed fair.

Why storing pirated copies was not

The same courts split a crucial hair: acquiring and retaining pirated copies from so-called shadow libraries was not protected, even if the eventual training use was. In the case that later settled for $1.5 billion covering nearly half a million works — the largest AI-copyright settlement to date, averaging about $3,000 per work before fees — the court found that downloading and maintaining a central library of more than seven million pirated books was not justified by training. There was, as the court put it, no get-out-of-jail-free card for obtaining infringing copies simply because the intended end use might be transformative. That exposure — magnified by statutory damages of up to $150,000 per work willfully infringed — is what ultimately pushed the settlement.

A parallel case against another platform reached a similar fair-use conclusion for training but went further, holding that training was fair use regardless of whether the books were pirated, while still leaving open claims about distributing pirated copies through torrent "seeding." The tension between those two approaches is now the central disagreement awaiting appellate review. Their final notice on seeding-related claims was still being briefed in early 2026 after newly produced log data suggested large-scale torrenting activity.

Meanwhile, a 2025 Delaware ruling in a non-generative AI context remains instructive. A court found that an AI-driven legal research tool that copied and used verbatim Westlaw headnotes to build a competing product was not fair use. The court stressed the use was commercial and non-transformative — same function, direct competitor — and that market harm weighed heavily against the defendant. The key distinction the court drew for your purposes: that tool did not generate new expressive content; it searched and surfaced. The case is now on appeal to the Third Circuit, which may clarify how much transformation is enough.

The output question is still live

If training doctrine is beginning to crystallize, output doctrine is not. A multidistrict litigation consolidated in the Southern District of New York in April 2025 — now covering twelve cases against a leading model provider — alleges that chat outputs reproduced detailed summaries and outlines of copyrighted books. In October 2025 the court denied a motion to dismiss, holding that some outputs alleged were sufficiently similar that a reasonable jury could find infringement, and left fair use for later. Discovery has since expanded to tens of millions of output logs, with orders in January and March 2026 compelling production of 20 million, then an additional 78 million and 10 million logs.

On the image side, a pending Central District of California case alleges an image-generation service copied protected characters to train its model and continues to produce derivative images of iconic franchises. The plaintiffs seek injunctive relief and willful-infringement damages. Just as notably, the same entertainment group and a major AI lab announced in late 2025 a three-year, $1 billion licensing and investment agreement allowing short AI videos featuring more than 200 characters — costumes, props, vehicles and all — with explicit commitments to creator rights.

That deal matters to your fair-use analysis even if you never license a character. Courts now have concrete evidence of an active, remunerative licensing market for AI training and generation. Where such a market exists, unlicensed uses that impair it weigh against fair use. For a small business, the inference runs the other way too: choosing a model that is licensed for the types of content you generate puts you on the safer side of that same factor.

What This Means for the Content You Publish

Translate those doctrines to three everyday scenarios.

If you generate a blog post, product description, or logo draft and publish it as-is, you almost certainly hold no copyright in that specific expression. A competitor who copies it verbatim is not infringing you, because you have no protected interest to infringe. Registration will be refused if you claim sole AI authorship.

Where protection does attach is in the human layer you add:

  • Curation and compilation: You generate 20 headlines and select, sequence, and edit five into a campaign. The selection and arrangement can be original.
  • Substantial revision: You use AI for a first draft and then rewrite for voice, add original examples, verify facts, and restructure. The revised expression is yours.
  • Hybrid assembly: You combine AI-generated elements with human photography, illustrations, or data visualizations into a new composite.

The Copyright Office's current guidance asks applicants to disclose AI-generated material and to claim only human-authored contributions. Explain what you did — "human-authored text with AI-assisted research" is more defensible than silence.

2. AI can still infringe when it reproduces too much of someone else's work

Even if training is fair use, an output that is substantially similar to a protected work can infringe. Two risk patterns showed up repeatedly in 2026 filings:

  • Near-verbatim text recall. Summaries or excerpts that track a book chapter's structure and phrasing too closely, especially when the model was prompted to reproduce it.
  • Character and style replication. Images that are instantly recognizable as a specific franchise's characters, even without naming them in the prompt, where the model has learned to render those features.

Your exposure as a user is not theoretical. If you publish an AI-generated image of a mouse-eared mascot that reads as a famous cartoon character, or a "write in the style of" article that lifts distinctive passages, the infringement claim runs against the publisher — you — not just the model provider.

Mitigation is practical, not legalistic:

  • Avoid prompts that ask for a specific living author's style, an existing brand's mascot, or a known character. Describe the function instead: "friendly retro robot mascot for a bike shop, original design, no reference to existing franchises."
  • Run a quick reverse-image search and a text-similarity check before publishing. If the output mirrors something you recognize, regenerate with a fresh prompt.
  • Prefer tools that filter for known copyrighted characters and that offer to block verbatim regurgitation. Confirm those safeguards are on.

3. Licensed vs. unlicensed models now create different paper trails

The 2026 licensing market changes procurement, not just doctrine. Vendors that can show licenses for training data or for character likenesses give you two advantages: lower downstream infringement risk and a credible indemnity story if a claim arrives.

When you evaluate tools, ask for specifics, not marketing. "Trained on licensed data" should come with what was licensed, from whom, and whether your use case is covered. A deal that covers video generation with 200 characters does not automatically cover your AI-generated picture book for sale. Get the scope in writing.

The Five Mistakes That Keep Showing Up

Across dispute logs that became public and Copyright Office refusals, the same missteps recur for small teams:

  1. Assuming AI output is automatically owned. You file no registration, keep no edit history, and later discover you have no enforceable right when a larger competitor copies your best-performing landing page.

  2. Publishing recognizable characters or brand look-alikes. A coffee shop's AI-generated mural featuring look-alike superheroes invited a cease-and-desist that cost more than commissioning original art.

  3. Prompting for "in the style of" a living creator. An agency delivered an AI-written newsletter that echoed a well-known author's distinctive voice and examples, inviting both an infringement concern and a client-relations problem.

  4. Keeping no records. No saved prompts, model versions, or edit trails. When a client asks "where did this image come from?" you have no answer that survives an audit or a dispute.

  5. Ignoring the tool's terms on ownership and indemnity. Some consumer-grade tools grant you ownership of outputs; others retain broad licenses to reuse your prompts and outputs, or offer no indemnity for third-party claims. The cheapest monthly plan is often the thinnest legally.

A Practical Checklist Before You Publish AI-Assisted Work

Use this six-step pass on every external asset — it takes minutes once it is habit:

  1. Add verifiable human authorship. Edit substantively and keep a tracked-changes version. Aim for at least a paragraph-level rewrite on text and layered edits on images, not just a spellcheck.

  2. Run similarity checks. For text, search distinctive sentences in quotes; for images, reverse-image search. If you find a near match, treat it as a redraft signal, not a coincidence.

  3. Save the provenance. Log the tool, model version, date, prompt, seed or settings, and the human edits in a shared folder. A one-line Beancount-style note next to the asset's cost entry works: "2026-08-10 generated 'summer sale hero' v4, prompt X, edited by J.S., approved."

  4. Choose tools with indemnity for your use case. Prefer business-tier plans that explicitly assign output ownership to you and offer IP indemnification where the output is used as directed. Keep the terms sheet with the subscription receipt.

  5. Disclose where it matters. Client-facing work, especially ghostwritten thought leadership or images presented as original illustration, should note AI assistance in your contract or delivery note. Disclosure now prevents a dispute letter later.

  6. Register thoughtfully. If you register a work containing AI-generated material, describe the human contribution and disclaim the AI-generated portions as the Copyright Office requires. Over-claiming invites refusal or later invalidation.

How to Bookkeep AI Content Costs So They Actually Help at Tax Time

The copyright question is entangled with a bookkeeping question. Whether you expense, capitalize, or simply track AI spend affects both your tax position and your ability to show prudent procurement if ownership is ever challenged.

Give AI its own chart. Most small businesses bury AI under "Software" or "Marketing." Break it out:

  • Expenses:Software:AI:Subscriptions — monthly seat fees, API usage, per-token charges
  • Expenses:Software:AI:Licensed-Content — fees for models or datasets licensed for commercial generation, including character or stock-image licenses
  • Expenses:Marketing:Content-Creation — human editing, fact-checking, and design time applied to AI drafts
  • Assets:Prepaid:Licenses — annual licenses paid upfront and amortized

This separation answers two questions quickly: how much are you spending on generation versus human finishing, and which tools carry licensed data worth renewing.

Track usage-based billing as it happens. Many generative tools bill per 1,000 tokens, per image, or per minute of audio. Reconcile the vendor's usage dashboard to your bank settlement weekly, not monthly. A simple weekly entry that records tokens or images, cost, and project — "8,200 tokens, $12.30, August campaign" — catches runaway experiments early and creates the contemporaneous record that supports both cost allocation and, if needed, a showing of good-faith use of licensed systems.

Allocate, don't lump. If one model produces blog posts, product images, and internal documentation, allocate cost by output. That allocation matters when you later need to know the true cost per publishable asset or to justify a write-off if an AI-generated campaign is pulled for rights review. Plain-text accounting shines here: a single transaction can split one payout across three expense accounts with clear narration.

Keep terms with receipts. Attach the vendor's IP-ownership and indemnity clause — a PDF or even a permalink — to the transaction metadata. When you later reconcile or your accountant asks why one tool costs three times another, the answer is in the ledger, not in someone's inbox.

For a deeper setup, see the plain-text account hierarchy examples in the Beancount documentation and visualize category spend over time with Fava to spot creep before it becomes a budget line you cannot explain.

What to Ask Your AI Vendor Before You Renew

Copy these four questions into your next renewal review. The answers belong in your files as much as the invoice does.

  • What data was the model trained on for my use case, and what is licensed? Vague "publicly available data" is not an answer. You want the categories and the license type.

  • Do you remove or preserve copyright management information? Removal that enables infringement has triggered separate claims under the Digital Millennium Copyright Act. Preservation is a positive signal.

  • What IP indemnification do you provide for outputs used as directed, and what voids it? Note exclusions for prompts that request a living author's style or a known character.

  • What filters prevent verbatim or near-verbatim outputs? Ask how the system blocks recall of protected text and character likenesses, and whether you can enable stricter settings for commercial publishing.

If a vendor cannot answer in writing, treat the price as carrying an unpriced risk premium.

Keep Your Content Useful and Defensible

You adopted generative AI to move faster without hiring faster. The businesses that keep that advantage in 2026 are not those that avoid the tools, but those that pair them with a lightweight discipline: add human craft you can point to, check before you publish, log what you used, and pay for licensed models where the stakes are highest. Copyright still rewards human expression, training doctrine is stabilizing in a way that favors transformative, licensed systems, and output liability remains where compliance is most controllable — at the moment you choose to publish.

That same discipline belongs in your books. Clear AI spend tracking turns a tangle of micro-charges into a deductible, auditable history of how you built your marketing engine. It also makes the next tool comparison honest, because you know what you actually paid per finished asset, not just per seat.

Simplify Your Financial Management

As you build a library of AI-assisted content, keeping a precise record of tooling costs, licensing fees, and human editing time is what lets you price, protect, and defend that work. Beancount.io gives you plain-text accounting that is transparent, version-controlled, and AI-ready — so your financial records are as auditable as your creative ones. Get started for free and see why technical teams prefer accounting they can read, diff, and own.

Share this article