MEMECMO: turning GEO from a slogan into something you can measure
How an AI product actually grew: the multiplication that gives it commercial value, three corrections we had to make to our own metrics, false positives in the judging layer, how to train "AI workers", and why we go deep in only a few industries.
Updated 2026-09-22
The short version
- MEMECMO is our own GEO platform: it measures whether a brand gets named, how it is described, and who outranks it in the answers of ChatGPT, Claude, Gemini, Perplexity and Google AI Overview.
- The hard part was never calling a model. It was making the measurement itself defensible: prompts that never name the brand, every rate printed with its denominator, and a judging layer that is itself audited.
- The value formula we settled on is a multiplication: base intelligence (API) × engineering delivery (Vibe Coding) × vertical functional capital (domain know-how). If any factor approaches zero, so does the result — which is why a general-purpose model cannot simply swallow an industry.
- You can check us: the demo brand’s entire scan on the Hong Kong GEO platform is public — 676 cells, 668 usable answers, every rate with its numerator and denominator.
Why we built it
By 2026, asking an AI before deciding is ordinary. Brands discovered something awkward: they rank respectably in search, and yet they simply do not appear in AI answers — not ranked low, absent.
A wave of "GEO" tools appeared, most of them working the same way: take the brand’s website and some SEO keywords, generate questions that contain the brand name, ask the models daily, draw a flattering line. But if the question already names the brand, of course the model mentions it. That measurement confirms itself, and any optimisation a client bases on it cannot be verified.
So we put the product’s first principle in the measurement: make "did the AI bring you up unprompted" measurable, reproducible and open to challenge, and only then talk about optimising it. That decision shaped all of MEMECMO, and it is why the Hong Kong GEO platform can publish an entire scan.
The value formula, and why it multiplies
Commercial value = base intelligence (API) × engineering delivery (Vibe Coding) × vertical functional capital (domain know-how)
Base intelligence is purchasable: a handful of labs sell APIs that get cheaper every year and better every month. Necessary — and available to everyone, so it is nobody’s moat.
Engineering delivery is the ability to turn that intelligence into something a business person can actually use. We call it Vibe Coding: hackathon pace on a low-code base, a working minimum viable product in weeks, not a system delivered in six months that nobody opens.
Vertical functional capital is the judgement a firm has accumulated: what counts as a real lead, which compliance wording must never be touched, what will annoy a regulator, how buyers in a category actually phrase a question. No public corpus contains it, because it was never written down.
They multiply because they do not substitute for one another. API without delivery is a demo. Delivery without domain judgement is a generic tool the client abandons in a fortnight. Domain judgement without delivery stays trapped in people’s heads, where it cannot be called, copied or audited.
Functional capital, reconsidered
- Functional capital
- The industry judgement and working method a company has accumulated but never structured. It lives in senior staff’s experience, past case files and the unwritten standards of internal review — not in any system.
Before generative AI, functional capital was undervalued: it spread only person to person, slowly, but nobody could take it either. Generative AI reversed that. General models turned general capability into a commodity, so what remains distinctive about a company is very nearly only its own judgement.
This is why we do not sell general-purpose systems. We take what the client already has, structure it, and turn it into assets an agent can call — reverse-engineering functional capital: leave the workflow alone, take cognitive slices from archives, case files and content, build a knowledge graph that is theirs, then ship a usable interface at pace.
In MEMECMO this became the brand profile: a numbered list of facts, every one of them already public on the client’s own channels. It is not marketing copy; it is the baseline for judging whether an AI said something false. The richer the profile, the more the error alerts are worth — and a thin profile makes the accuracy metric nearly meaningless. We learned that the hard way.
Our own functional capital: why this team could build it
The concept has to apply to us too. NeuronSpark is not only the builder of MEMECMO but a long-term partner in it, and holds equity — worth disclosing before you weigh anything below. We were willing to tie ourselves in that closely because the work sits exactly where two kinds of capital meet.
The first is engineering delivery. The rules above — resumable long jobs, narrow evidence-bound judging, never publishing a scan that lost a tenth of its cells — are not method-deck prose; they are running code. Handing a 676-cell scan to a scheduler and trusting the result is an engineering achievement, not a model one.
The second is a communications background. Our founder spent over twenty years in digital media and communication studies, and co-founded a high-reach finance publication in China. That history set our basic reading of GEO: AI visibility is a communications problem, not a tagging problem. How buyers phrase a question, which formulations get repeated, how media and encyclopaedias become the source an AI leans on — none of that grows out of an SEO tool. It comes from people who have done communications. Prompts that never name the brand, and three metrics kept strictly apart, are direct consequences of that reading.
We chose our closest partner on the same criterion. MEMECMO’s founder brings two decades of enterprise consulting and is an investor, with deep first-hand relationships inside specific industries. That is why our prompt sets can be built from what buyers actually ask rather than guessed from a keyword tool, and why, at delivery time, we can find people in the industry willing to have their public facts checked. The team also carries public-relations experience from multinational technology companies, and that experience shaped a delivery choice: rather than build one comprehensive system over six months, move in short steps and get a single business line working first.
Stack those together and MEMECMO is not "a team that wrapped an API into a GEO tool". It was built with industry sediment: agents polling and dividing the work, every scan strengthening the prompt set, the alias table and the golden sample set, and a judging standard that can itself be tested and replaced. Judgement therefore compounds — which is precisely the factor in the opening formula that cannot be bought.
This is also the line between us and the two kinds of GEO tool on the market. One is a pure algorithm still scoring SEO tags under a new name. The other wraps a large lab’s prompts and ships; all of its judgement is on loan from someone else’s model and drifts the moment that model changes. What we accumulate is our own data assets and judging standards — they do not expire when a model does.
The direction from here is to train a dedicated small model on that sediment and cut the cost of the judging layer by another order of magnitude. That step has not been taken; it is stated here as a roadmap, not a present capability.
What GEO actually is: three metrics that must stay apart
- GEO (generative engine optimisation)
- The work of getting a brand mentioned and described correctly in generative AI answers: measure the present state, find the gaps, change the public information an AI can read, then measure the change.
Early on we made the classic mistake of blurring three different numbers. We then forced them apart and required every rate to be displayed with its numerator and denominator.
- Presence = answers naming the brand ÷ usable answers. Computed per brand; the brands do not add up to 100%.
- Share of voice = that brand’s mentions ÷ all brands’ mentions. This one does add up to 100%.
- Citation = how often the sources an answer links to include the brand’s own domain. It measures what the AI is basing its answer on.
Tools routinely conflate them, and sometimes invert them. In August 2026 one of our own glossary documents had presence and share of voice the wrong way round; we only caught it when re-checking against real data. The cost of that class of error is not embarrassment — it is that a client optimises the wrong thing. The rule since: every rate appears with its denominator, in the product, in the report and in the email.
What went wrong, and the rule each failure left behind
Every item below happened in production; none is a hypothetical risk.
- False positives in judging. We first had a model judge six answers at once and flag "what is wrong here". One scan produced 61 flagged errors. Re-checking each claim individually — every claim had to name the numbered profile line it contradicted — left exactly 1 standing. Rule: judging must be narrow and evidence-bound, never open-ended.
- Brand names fracture. The same company appears under many spellings; one scan produced 232 name variants. We merge them at read time and rank only within the 12 tracked brands, or a provider mentioned once with one kind word outranks the client on sentiment. Rule: a ranking needs an explicit comparison set.
- Generic words poison the data. A phrase like "Hong Kong telecom operator" contains a client’s own name, and naive matching counts it as a mention. Rule: exclude generic phrases explicitly and review the alias table by hand.
- Engines get retired. Mid-run, a provider withdrew one of the models we were using. Rule: re-verify the model list regularly and state in the report which models that period actually used.
- Credit runs out. Once, an exhausted payment balance made the platform mark hundreds of cells as "failed" mid-scan, and the dashboard showed a brand at almost zero — wrong data that looked entirely real. Rule: account-level failures do not count as failed answers; the cells stay pending, and a scan that loses more than a tenth of its cells is never published — the dashboard keeps showing the previous period.
All of these rules are now code in the Hong Kong GEO platform. They look like housekeeping, and they are exactly what decides whether a client can trust your numbers.
Delivery: from six months to a few weeks
Enterprise software takes six months because it assumes the requirements can be stated before the work starts. AI products break that assumption: you learn where the model’s competence ends by running it, and the client can only tell you what they want once they see an interface.
So delivery is cut into phases you can stop at: weeks 1–2, find the target and hand over a structured assessment; weeks 3–4, ship the first vertical agent interface and run it in the client’s real environment; months 2–3, extend sideways and join the agents into one working flow. The client can stop at the end of any phase holding something real.
One engineering discipline follows from this: long jobs must be resumable. A full GEO scan is 676 external calls and a cloud function may run for five minutes, so the scan is built to stop anywhere and continue from there rather than start over. That discipline is what made it safe to hand scanning to a scheduler.
How to train an "AI worker"
Treating an agent as a colleague rather than a function is the single most useful habit we have. Colleagues need a job description, a standard of judgement, review, and a qualifying test. So do agents.
- Fix the job description. One agent, one job, fixed inputs and outputs. "While you are at it, check for anything else wrong" is where incidents begin.
- Make the standard nameable. Do not ask "is this answer correct". Ask "which numbered line of the profile does this sentence contradict". The first gets you an opinion; the second gets you evidence.
- Qualify before hiring. We keep a blind, human-labelled golden set. Any change of judging model must first pass an agreement check against it. Models turn over constantly; the standard must not drift with them.
- Send low confidence back to a human. When a judgement carries a probability, anything below the threshold goes to review instead of being decided by the system.
- Keep the trail. Every judgement must be traceable to the original answer and its evidence, or the only thing you can tell a client is "the model said so".
The same applies on the client side. The interfaces we ship have no learning curve — staff call the analysis in plain language — but we insist the client names one person to confirm the brand profile and the prompt set, because the functional capital is theirs, not ours.
Why we go deep in only a few industries
We have turned down markets that look large. Three criteria, all required.
- Buyers really do ask AI. For telecom, home broadband or cross-border consumer goods, buyers ask an AI which one is worth it. In industries where procurement runs on relationships and tenders, AI visibility does not move a deal.
- The question can omit the brand. The industry must have natural category questions ("which business broadband should a Hong Kong SME choose"), or the measurement only confirms itself.
- The facts are public and checkable. Profile facts must already be published on the client’s own channels before we dare judge an AI wrong against them. Where facts are not public, accuracy cannot be measured at all.
One boundary is self-imposed: we work on outbound visibility only, not on AI engine optimisation inside mainland China. The reason is not technical. This method rests on publicly checkable facts and measurement a third party can reproduce, and those premises hold to different degrees in different information environments. Rather than compromise between two sets of rules, we do one thing well enough to be verified.
What you can check right now
A method is persuasive when its own data is open, not when it is well phrased. Three things anyone can inspect:
One honest footnote on accuracy: that scan scored 99.8% (600 / 601), but the brand profile held only 13 facts at the time. "Accurate" largely meant "there was little to check against". That is precisely why the profiling phase is billed separately — the density of the profile decides whether the metric means anything.
If you are building in the AI era
- Get the measurement right before optimising. Optimisation you cannot measure cannot be accepted, and the client will eventually ask for proof.
- Print the denominator with every rate. It prevents half the misunderstandings and almost all of the self-deception.
- Audit the auditor. Who checks the checker is a question that must have an answer.
- Cut delivery into phases a client can stop at. You can promise something specific precisely because they can walk away.
- The moat is on the client’s side, not the model’s. Your job is to dig it out and structure it, not to paper over it with general capability.
Frequently asked
- What is NeuronSpark’s relationship with MEMECMO?
- NeuronSpark built MEMECMO, works on it as a long-term partner, and holds equity in it. The tie is that close because the two sides’ capital is complementary: NeuronSpark brings engineering delivery and a communications reading of visibility, while MEMECMO’s founder brings two decades of enterprise consulting and first-hand relationships inside specific industries. We disclose the relationship here so readers can weigh the conclusions accordingly.
- How is GEO different from SEO?
- SEO optimises the ordering of results. GEO is about whether a generative AI brings you up at all when answering a question, how it describes you, and whose material it cites. One produces a link; the other produces a paragraph the user takes as the conclusion.
- What makes an AI-visibility measurement credible?
- Three conditions, all required: the prompts never name the brand (otherwise it only confirms itself); the prompt set is frozen and reused period after period; every rate is shown with its numerator and denominator. We also ask each question in several phrasings across several engines to damp the randomness of any single answer.
- What exactly is functional capital?
- The industry judgement a company has accumulated but never structured: what counts as a real lead, which wording is non-compliant, how buyers actually ask. It sits in no system, so it cannot be learned from public text — which is what makes it more valuable in the AI era, not less.
- Why is commercial value a multiplication rather than a sum?
- Because the three factors do not substitute. API without delivery stays a demo; delivery without domain judgement is a generic tool; domain judgement without delivery stays locked in people’s heads. Any factor near zero takes the product with it.
- How long does a usable AI product take?
- On our phased approach: weeks 1–2 for the diagnosis and its assessment report, weeks 3–4 for the first interface running in the real environment, months 2–3 to extend it into a working flow. Each phase ends in something you can accept or reject.
- Why not work on AI engines inside mainland China?
- The method depends on publicly checkable facts and measurement a third party can reproduce. We chose to work on outbound visibility only, and to make that one thing verifiable, rather than compromise between two different information environments.