TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Generative Mechanics · AGE 15—17 · FACTUAL · EVOLVING
Inputs beyond text
A tool that cannot take a photograph cannot look at your site.
The idea
Until recently, the chat tools most people met took only typed words. Today many tools also take a photograph, a PDF, a screenshot, a voice recording, and sometimes a short video. Each kind of input goes through its own front end first — words through a tokeniser, pictures through a vision encoder, sound through an audio encoder, a PDF through text extraction or page images — and only then does the model work on the numbers. What a tool can take in decides what you can ask of it. A tool that cannot take a photograph cannot look at your site.
Why it matters
Before you ask "what can it do?", ask "what can I give it?" The second question is quicker to answer, and it tells you more.
See it in the studio
You photograph cracked plaster on a site wall and want to ask about it. One app takes the photo and looks at it. Another only takes words, so you have to describe the crack. Your description is already a guess about what matters.
Watch for this
Assuming that because a tool accepts a file, it reads the file the way you do. Taking the file in is one thing. How well it reads the file is a separate question, and you have to test it.
Try it
Before you upload anything: it must be yours, or cleared for this use — see FOUR RISKS, the map's gate before any upload. Then take one AI tool you use. List what it can take in: typed text, a photo, a PDF, a voice note, a screenshot, a video. Tick what works. Give it one photo of a building and ask what it sees. Check the answer against the photo. Now give it a photo with a Kannada signboard, or a drawing with Kannada labels, and ask what it reads. Ask the same of an English sign. Note what it got right in each script. That difference is part of what "can take in" means.
Prove it
Name three kinds of input a tool can take besides typed words, and say one thing a tool could not do if it only took words.
How it works
Words are cut into tokens. Pictures are cut into small patches by a vision encoder. Sound is cut into short slices by an audio encoder. Each front end is different, and each produces a list of numbers; many current models then put all of them into one stream and reason over the mix (ONE STREAM on this strand). Multimodal systems existed in research long before the chat products; what changed was that ordinary users got them. The list of accepted inputs changes with every product release. That is why this card is marked EVOLVING: the idea holds, but the list does not.
What this idea builds on
- A starting idea.
What this idea opens up
Sources
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.