TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Generative Mechanics · AGE 15—17 · FACTUAL · EVOLVING

Inputs beyond text

A tool that cannot take a photograph cannot look at your site.

The idea

Until recently, the chat tools most people met took only typed words. Today many tools also take a photograph, a PDF, a screenshot, a voice recording, and sometimes a short video. Each kind of input goes through its own front end first — words through a tokeniser, pictures through a vision encoder, sound through an audio encoder, a PDF through text extraction or page images — and only then does the model work on the numbers. What a tool can take in decides what you can ask of it. A tool that cannot take a photograph cannot look at your site.

Why it matters

Before you ask "what can it do?", ask "what can I give it?" The second question is quicker to answer, and it tells you more.

See it in the studio

You photograph cracked plaster on a site wall and want to ask about it. One app takes the photo and looks at it. Another only takes words, so you have to describe the crack. Your description is already a guess about what matters.

Watch for this

Assuming that because a tool accepts a file, it reads the file the way you do. Taking the file in is one thing. How well it reads the file is a separate question, and you have to test it.

Try it

Before you upload anything: it must be yours, or cleared for this use — see FOUR RISKS, the map's gate before any upload. Then take one AI tool you use. List what it can take in: typed text, a photo, a PDF, a voice note, a screenshot, a video. Tick what works. Give it one photo of a building and ask what it sees. Check the answer against the photo. Now give it a photo with a Kannada signboard, or a drawing with Kannada labels, and ask what it reads. Ask the same of an English sign. Note what it got right in each script. That difference is part of what "can take in" means.

Prove it

Name three kinds of input a tool can take besides typed words, and say one thing a tool could not do if it only took words.

How it works

Words are cut into tokens. Pictures are cut into small patches by a vision encoder. Sound is cut into short slices by an audio encoder. Each front end is different, and each produces a list of numbers; many current models then put all of them into one stream and reason over the mix (ONE STREAM on this strand). Multimodal systems existed in research long before the chat products; what changed was that ordinary users got them. The list of accepted inputs changes with every product release. That is why this card is marked EVOLVING: the idea holds, but the list does not.

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON