A working product with ongoing research and development
My Words exists because of a harder problem than translating a sentence: holding a conversation together across a language boundary when the stakes are somebody's housing, safety or health, and doing it in a way that tells the truth about its own uncertainty. That problem is not finished, so the research did not stop when the product shipped.
Where it came from
The work began as Project Keystone, research by INT6 Ltd into real-time bilingual conversation mediation using language models. INT6 built the proof-of-concept application, Turnstone, which was then developed into a working system and tested in real settings.
The application has since been moved into C2 Discovery Labs CIC so that it is held for community and societal benefit. My Words is that same tested system, rebranded and carried forward, and the development continues under it.
Academic validation
Independent validation is being arranged with a UK university partner. The intention is to work with postgraduate students and academic research labs on two things at once: improving the product, and checking that it does what this site says it does.
That work is not finished and nothing on this site depends on it. When there are results, including any that go against us, they will be published here with the method attached. Until then, treat the claims on this site as ours rather than as independently checked.
What the project is actually trying to find out
These are the questions the work is organised around, and the ones we would put in front of a research partner tomorrow. None of them is settled.
Can a machine hold context the way an interpreter does?
A human interpreter carries the room, the history and the register without being asked. How much of that can be made explicit as structured context, and what is lost in making it explicit?
Can a system usefully know when it is wrong?
Fidelity scoring is a model judging a model. Where does that correlate with real error, and where does it fail in the same direction as the thing it is checking?
Does a clarifying question help or interrupt?
Asking beats guessing, in principle. With a distressed person in front of you, how many times can you interrupt before you have damaged the thing you were protecting?
How should uncertainty be shown to a busy professional?
A marker that is ignored is no better than no marker. A marker that stops the conversation every turn is worse. This is an interface research question as much as a modelling one.
What changes when the person holds the device?
Reading in your own hands, in your own language, is a different act from reading across a desk. Whether it changes what people disclose is a question for measurement, and we have not measured it yet.
Where does the boundary really sit?
We assert that this is for early conversations and not for statutory ones. That boundary should be tested against what actually happens in services, not defended because we wrote it down.
How the work is evaluated
The system produces its own evidence as a by-product of running: every turn carries a source, a translation, a fidelity score, any revision, and any clarifying question. That makes several things measurable without instrumenting anything extra.
- How often a translation is revised before delivery, and by how much the score moves.
- How often the mediator asks instead of answering, and what the conversation does next.
- How scores distribute by language, so that uneven coverage shows up instead of disappearing into an average.
- Where workers overrode, repeated themselves, or abandoned the exchange.
What the instrumentation cannot tell us is whether the conversation was any good. That needs the service and, where it can be done properly, the person on the other side of it.
How we hold ourselves to it
Test before live. New behaviour is exercised against recorded and synthetic conversations before it meets anyone's real one.
Measure what was measured. No headline accuracy figure standing in for a distribution we can actually show.
Publish the method. What we learn about doing this safely is written up for other people to reuse, including the parts that did not work.
Negative results count. A feature that made conversations worse is a finding, and gets reported as one.
What is built and running
This is a working application, not a demonstration. In broad terms it currently has:
Live turn-by-turn mediation
Streaming translation with conversation history, structured facts and domain context in scope for every turn.
Fidelity scoring and self-revision
Per-turn scores, below-threshold revision before delivery, and clarification back to the speaker where the source is ambiguous.
Two conversation surfaces
A shared-screen mode and a paired mode where the second person joins from their own device over a live connection.
A localised interface
Interface strings translated into the seeded languages and cached, so the surface itself is not English-only.
Accounts, groups and audit
Organisation-level grouping and permissions, passkey and password authentication with two-factor codes, and an audit log across conversations and administration.
Replayable conversations
Recorded exchanges can be replayed against a new build, so a change is tested on real dialogue and not on something improvised for the demo.
An ordinary web stack
A PHP and MariaDB application served over HTTPS. Deliberately unexotic, so it can be reasoned about, audited, and hosted somewhere other than where it sits today if that becomes the right answer.
Work on this with us
We are looking for postgraduate students and research labs with a question in this space, for services willing to have their use studied, and for funders backing evidence-led work on language as a barrier to getting help.