¿Puede un programa saber si un texto lo escribió una máquina?
No.
Ninguno puede. Ni este, ni los de pago, ni el que le enseñaron en una formación.
Si alguien le dice lo contrario, le está vendiendo algo.
Esta página explica, sin una sola palabra técnica, qué hace entonces esta herramienta,
cuánto se equivoca —con el número medido, que casi nadie publica— y qué puede usted
hacer el lunes con todo eso.
Entonces, ¿qué hace?
Tres cosas distintas. Y lo importante no son las tres: es que dos son hechos y una
es una opinión, y confundirlas es de donde salen las injusticias.
Opinión sobre la prosa
La puntuación. Señala palabras y giros que la escritura de máquina usa mucho, y le
enseña cada uno subrayado en el texto. Es un juicio sobre cómo suena. Se puede discutir,
y a veces se equivoca.
Hecho
Caracteres que no se teclean. Las herramientas que «humanizan» texto dejan letras
de otro alfabeto que se ven idénticas. Le da el código, la línea y la columna. Está o no está.
Hecho
La bibliografía contra sí misma. Una fuente citada que no aparece en su propia lista;
el mismo identificador en dos artículos distintos; un año que aún no ha llegado. Sin internet:
el documento se contradice solo.
Un porcentaje se discute en una reunión durante media hora. Que el mismo identificador
esté en dos referencias distintas se resuelve con una pregunta: «¿me manda el artículo?».
Por eso la herramienta enseña las dos cosas por separado y nunca suma la segunda a la primera.
¿Cuánto se equivoca?
Esta es la pregunta que ningún detector contesta, y es la única que importa cuando hay
una persona delante.
Lo medimos así: noventa textos publicados antes de que existieran los modelos
generativos. No los juzgó nadie: sabemos que son humanos por su fecha. Después los
pasamos por la herramienta y contamos a cuántos habría acusado.
Ninguno de los noventa fue marcado. Pero cero de noventa no significa que la tasa
sea cero: con esa muestra, lo máximo que se puede afirmar honestamente es que se
equivoca menos del 4,1 % de las veces. Esa franja ámbar es la diferencia
entre una medición y una promesa.
Y no es el mismo número para todos
Repartir la cifra global a todo el mundo sería cómodo y sería falso. El corpus tiene
sesenta y cinco textos en inglés y veinticinco en español, así que lo que se puede
afirmar sobre cada idioma es distinto — y en la herramienta cada lector ve el suyo,
no el promedio.
Menos textos en español significa menos certeza, no más errores. La barra es más larga
porque sabemos menos, y decirlo es más útil que esconderlo detrás del promedio.
Lo que no le dice
Cuánta escritura de máquina se le escapa. No está medido, y es a propósito:
para medirlo habría que reunir textos generados, y esa colección solo representa los modelos
que estaban de moda ese mes. Una herramienta que no marca nada tiene una tasa de error perfecta.
Los noventa textos son artículos publicados, no trabajos de estudiantes. Un
ensayo de primero de carrera no se parece a un artículo revisado por pares, y esa diferencia
no está medida aquí.
Y una que conviene saber antes de mirar ninguna puntuación: la señal que más salta en escritura
humana es el ritmo uniforme de las frases — aparece en el 27,8 % de los textos
humanos que medimos. Si la evidencia de un trabajo se apoya sobre todo en eso, vale menos.
La herramienta lo dice en cada informe, por su nombre.
¿Por qué fiarse de este y no de otro?
No por la tecnología. Por esto: el 10 de agosto de 2026 esta herramienta acusaba
en falso, y lo publicamos.
Un trabajo con la bibliografía en orden, que solo listaba una lectura adicional sin citarla,
se anunciaba en el informe como «1 contradicción de fuentes», encima
de la frase «el documento se contradice a sí mismo». En la página que un profesor imprime y
lleva a un comité.
Lo peor no fue el fallo. Fue que el propio código sabía la diferencia: hay un
campo que distingue una contradicción real de un simple desorden, con un comentario que explica
que la gente lista lectura adicional con toda legitimidad. La línea que escribía el titular
nunca le preguntó. Estuvo cinco días publicado y no lo encontró ningún usuario.
Está arreglado, está contado y el historial es público. Pregúntele a cualquier detector
de pago cuándo fue la última vez que acusó en falso. Esa respuesta, y no un porcentaje,
es lo que debería decidir de quién se fía.
¿Y el lunes qué hago?
Nada de esto sirve si no se traduce en algo que decir en un aula o en una reunión. Eso está
escrito, en castellano, y no lleva software dentro:
La hoja del estudiante lleva impresa la cifra medida para su idioma, no la global.
Si el trabajo está en español, el número que va al comité es el 13,3 %.
Se ejecuta entera en su navegador. El texto no se sube a ningún sitio — tampoco a nosotros.
Esta página tampoco pide nada a ningún servidor: ni una fuente, ni una imagen, ni una analítica.
Can software tell whether a machine wrote this?
No.
None of them can. Not this one, not the paid ones, not the one demonstrated at a training day.
Anyone who tells you otherwise is selling something.
This page explains, without a single technical word, what this tool does instead, how often it is
wrong — with the measured number, which almost nobody publishes — and what you can do with any of
it on Monday morning.
So what does it do?
Three separate things. What matters is not the three: it is that two of them are facts and
one is an opinion, and confusing those is where the injustices come from.
An opinion about prose
The score. It names words and turns of phrase that machine writing leans on, and shows
you each one underlined in the text. It is a judgement about how the writing sounds. It can be
argued with, and sometimes it is wrong.
A fact
Characters nobody types. Tools that "humanise" text leave behind letters from another
alphabet that look identical. It gives you the codepoint, the line and the column. Either it is
there or it is not.
A fact
A bibliography against itself. A source cited in the text and missing from its own
list; the same identifier on two different papers; a year that has not happened. No internet
needed: the document contradicts itself.
A percentage can be argued about for half an hour in a meeting. The same identifier on two
different references is settled with one question: "can you send me the paper?" That is
why the tool shows the two apart, and never adds the second to the first.
How often is it wrong?
This is the question no detector answers, and the only one that matters when there is a person
on the other side of it.
Here is how it was measured: ninety texts published before generative models
existed. Nobody judged them human — their dates did. Then they went through the tool and
we counted how many it would have accused.
None of the ninety was flagged. But zero out of ninety is not a rate of zero:
with a sample that size, the most anyone can honestly claim is that it is wrong
less than 4.1% of the time. That amber band is the difference between a
measurement and a promise.
And it is not the same number for everybody
Handing the pooled figure to everyone would be convenient and false. The corpus holds sixty-five
English texts and twenty-five Spanish ones, so what can be claimed about each language differs —
and in the tool, each reader is shown their own, never the average.
Fewer Spanish texts means less certainty, not more mistakes. The bar is longer because
we know less, and saying so is more use than hiding it behind an average.
What it does not tell you
How much machine writing it misses. That is not measured, deliberately: measuring
it would mean collecting generated text, and any such collection represents whichever models were
fashionable that month. A tool that flags nothing has a perfect error rate.
The ninety texts are published articles, not student work. A first-year essay
does not read like a peer-reviewed paper, and that difference is not measured here.
And one worth knowing before you look at any score: the signal that fires most often on human
writing is uniform sentence rhythm — it appears in 27.8% of the human texts we
measured. If the evidence against a piece of work leans mostly on that, it is worth less. The
tool says so on every report, by name.
Why trust this one and not another?
Not because of the technology. Because of this: on 10 August 2026 this tool was accusing
people falsely, and we published that.
A piece of work with a perfectly good bibliography, which merely listed some further reading
without citing it, was announced in the report as "1 source contradiction",
above the sentence "the document disagrees with itself". On the page a teacher prints and carries
into a committee.
The bug was not the worst part. The code itself knew the difference: there is a
field that separates a real contradiction from mere untidiness, with a comment explaining that
people legitimately list further reading. The line printing the headline never asked it. It was
published for five days and no user found it.
It is fixed, it is written up, and the history is public. Ask any paid detector when it
last accused someone falsely. That answer, not a percentage, is what should decide who
you trust.
What do I do on Monday?
None of this is worth anything until it becomes something you can say in a classroom or a
meeting. That part is written, and there is no software in it:
The student sheet carries the figure measured for their language, not the pooled one.
If the work is in Spanish, the number that goes to the committee is 13.3%.
It runs entirely in your browser. The text is not uploaded anywhere — not even to us. This page
requests nothing from any server either: no font, no image, no analytics.