Choose AI tools by the job you need done, not by brand: each category has its own test for picking a tool and one main risk to watch.
- Categories outlast tools. New tools launch daily, so the atlas maps jobs rather than products.
- Check your data terms. For chat assistants, confirm whether your data may be used for training; do not assume.
- Reference images matter. For image tools, reference image support is now the real control surface.
- Watch for silent failure. Automation agents need logs and human checkpoints before anything irreversible.
- Verify the sources. Click through citations and skim the originals on anything that matters.
New tools launch daily; categories barely change. This atlas maps the territory by what you are trying to do — brand-agnostic on purpose, so it stays true after the next hundred launches.
Writing & thinking — chat assistants
Do: drafting, rewriting, summarizing, planning, tutoring, rubber-ducking. Choose by: context length for your documents, reasoning tier for hard problems, and whether your data may be used for training (check, do not assume). Watch for: confident fabrication — anything factual gets the five-question loop.
Watch for confident fabrication. Choose a chat assistant by context length, reasoning tier and data-training terms, and check anything factual.
Building — coding assistants & agents
Do: autocomplete on the light end; on the heavy end, agents that take a ticket, edit a repo, run tests, and open a PR. Choose by: how it handles your codebase's size, and the quality of its self-verification loop. Watch for: the review bottleneck — machines write faster than humans can check, and unchecked merges are how outages ship.
Review is the bottleneck. Coding agents write faster than humans can check, and unchecked merges are how outages ship.
Images — generation & editing
Do: text-to-image, instruction-based editing, style and character consistency via references, upscaling, background surgery. Choose by: whether it supports reference images (the real control surface now) and readable in-image text if you need it. Watch for: license terms on commercial use, and provenance metadata — keep it when you export.
Pick image tools that accept reference images. Check the licence terms for commercial use, and keep provenance metadata when you export.
Video & motion
Do: text/image-to-video clips, start-and-end-frame conditioning, increasingly with synchronized audio. Choose by: clip length, camera-language obedience, and cost per finished second — failed takes are the real price. Watch for: one dominant motion per shot; storyboard first.
Failed takes are the real cost. Judge video tools by clip length, camera obedience and cost per finished second, and storyboard one dominant motion per shot.
Voice & music
Do: narration, dubbing, cloning (with consent), full-song generation. Choose by: emotional range and language coverage. Watch for: consent and disclosure — cloned voices are the sharpest dual-use edge in the consumer toolbox, covered in Deepfake Self-Defence.
Knowing — answer engines & research
Do: search that answers in prose with citations; deep-research modes that read dozens of sources. Choose by: citation quality — can you actually click through and verify? Watch for: the summary flattening the disagreement between its own sources; skim the originals on anything that matters.
Citation quality decides it. Make sure you can click through and verify, because a summary can flatten disagreement between its sources.
Meetings & capture
Do: transcription, speaker labels, summaries, action items. Choose by: accuracy on your accents and jargon. Watch for: consent (recording laws are real) and where recordings are stored.
Automation — agents & workflows
Do: multi-step jobs across apps: triage the inbox, update the tracker, file the report. Choose by: connector breadth (see skills & tools) and how visibly it shows its work. Watch for: silent failure — insist on logs and human checkpoints before anything irreversible.
Insist on logs and human checkpoints. Agents can fail silently, so pick ones that show their work before anything irreversible happens.
Running your own — local runners
Do: download open-weight models and run them on your hardware, private by construction. Choose by: your VRAM, frankly. Watch for: the model license, which varies more than the marketing suggests.
Checking — evals & provenance
Do: test model output against your rubric; verify content credentials on media. Choose by: fit to your real cases. Watch for: nothing — this category is the watching. It is also the easiest one to skip, which is why this site keeps a record section on measurement and detection.
Part of the Stay Human record. Continue: finding content gaps →