¿Puede un programa saber si un texto lo escribió una máquina?

No.

Ninguno puede. Ni este, ni los de pago, ni el que le enseñaron en una formación. Si alguien le dice lo contrario, le está vendiendo algo.

Esta página explica, sin una sola palabra técnica, qué hace entonces esta herramienta, cuánto se equivoca —con el número medido, que casi nadie publica— y qué puede usted hacer el lunes con todo eso.

Entonces, ¿qué hace?

Tres cosas distintas. Y lo importante no son las tres: es que dos son hechos y una es una opinión, y confundirlas es de donde salen las injusticias.

Opinión sobre la prosa

La puntuación. Señala palabras y giros que la escritura de máquina usa mucho, y le enseña cada uno subrayado en el texto. Es un juicio sobre cómo suena. Se puede discutir, y a veces se equivoca.

Hecho

Caracteres que no se teclean. Las herramientas que «humanizan» texto dejan letras de otro alfabeto que se ven idénticas. Le da el código, la línea y la columna. Está o no está.

Hecho

La bibliografía contra sí misma. Una fuente citada que no aparece en su propia lista; el mismo identificador en dos artículos distintos; un año que aún no ha llegado. Sin internet: el documento se contradice solo.

Un porcentaje se discute en una reunión durante media hora. Que el mismo identificador esté en dos referencias distintas se resuelve con una pregunta: «¿me manda el artículo?». Por eso la herramienta enseña las dos cosas por separado y nunca suma la segunda a la primera.

¿Cuánto se equivoca?

Esta es la pregunta que ningún detector contesta, y es la única que importa cuando hay una persona delante.

Lo medimos así: 296 textos escritos antes de 2022 — artículos publicados, revisiones de Wikipedia y ensayos de clase. No los juzgó nadie: sabemos que son humanos por su fecha. Después los pasamos por la herramienta y contamos a cuántos habría acusado.

296 TEXTOS HUMANOS 2 marcados los dos rellenos, de doscientos noventa y seis LO QUE SE PUEDE AFIRMAR DE VERDAD 0,7 % 0 % 2,4 % 10 %
Dos de los 296 fueron marcados: el 0,7 %. Pero dos de 296 tampoco es exactamente la tasa: con esa muestra, lo máximo que se puede afirmar honestamente es que se equivoca menos del 2,4 % de las veces. Esa franja ámbar es la diferencia entre una medición y una promesa. No es cero, y ese es justamente el punto: hasta septiembre esta página enseñaba un cero, porque el corpus todavía no contenía a las personas con más probabilidades de ser acusadas en falso.

Y no es el mismo número para todos

Repartir la cifra global a todo el mundo sería cómodo y sería falso. El corpus tiene 271 textos en inglés y veinticinco en español, así que lo que se puede afirmar sobre cada idioma es distinto — y en la herramienta cada lector ve el suyo, no el promedio.

Inglés 271 textos 2,7 % Español 25 textos 13,3 % 0 % 5 % 10 % 15 %
Menos textos en español significa menos certeza, no más errores. La barra es más larga porque sabemos menos, y decirlo es más útil que esconderlo detrás del promedio.

Lo que no le dice

Cuánta escritura de máquina se le escapa. No está medido, y es a propósito: para medirlo habría que reunir textos generados, y esa colección solo representa los modelos que estaban de moda ese mes. Una herramienta que no marca nada tiene una tasa de error perfecta.

Sí incluye trabajos de estudiantes, y esa es la parte que más le importa a usted. De los 296 textos, 206 son ensayos de clase escritos por adultos que estaban aprendiendo inglés. Es exactamente el grupo del que se acusa a los detectores de equivocarse más, y aquí es el que peor sale: mide más alto que cualquier otro grupo en todas las columnas. A un umbral de 25 sobre 100 se marcaban nueve de ellos. Por eso la frontera está en 30 y no en 25 — se movió cuando supimos a quién estaba haciendo daño.

Lo que sigue sin medirse es el ensayo de estudiante en español: los veinticinco textos en español son revisiones de Wikipedia de una sola fuente. Ahí esta herramienta sabe menos, y lo dice en vez de prestarle a un idioma una cifra que midió en otro.

Y una que conviene saber antes de mirar ninguna puntuación: las señales que más saltan en escritura humana casi no dicen nada. El ritmo uniforme de las frases aparece en el 32,1 % de los textos humanos que medimos, y la palabra «just» en el 34,8 %. Si la evidencia de un trabajo se apoya sobre todo en eso, vale menos. La herramienta lo dice en cada informe, por su nombre.

¿Por qué fiarse de este y no de otro?

No por la tecnología. Por esto: el 10 de agosto de 2026 esta herramienta acusaba en falso, y lo publicamos.

Un trabajo con la bibliografía en orden, que solo listaba una lectura adicional sin citarla, se anunciaba en el informe como «1 contradicción de fuentes», encima de la frase «el documento se contradice a sí mismo». En la página que un profesor imprime y lleva a un comité.

Lo peor no fue el fallo. Fue que el propio código sabía la diferencia: hay un campo que distingue una contradicción real de un simple desorden, con un comentario que explica que la gente lista lectura adicional con toda legitimidad. La línea que escribía el titular nunca le preguntó. Estuvo cinco días publicado y no lo encontró ningún usuario.

Está arreglado, está contado y el historial es público. Pregúntele a cualquier detector de pago cuándo fue la última vez que acusó en falso. Esa respuesta, y no un porcentaje, es lo que debería decidir de quién se fía.

¿Y el lunes qué hago?

Nada de esto sirve si no se traduce en algo que decir en un aula o en una reunión. Eso está escrito, en castellano, y no lleva software dentro:

Texto para el programa de la asignatura Hoja de una página para el estudiante Procedimiento para un comité

La hoja del estudiante lleva impresa la cifra medida para su idioma, no la global. Si el trabajo está en español, el número que va al comité es el 13,3 %.

Probar la herramienta Ver el código

Se ejecuta entera en su navegador. El texto no se sube a ningún sitio — tampoco a nosotros. Esta página tampoco pide nada a ningún servidor: ni una fuente, ni una imagen, ni una analítica.

Can software tell whether a machine wrote this?

No.

None of them can. Not this one, not the paid ones, not the one demonstrated at a training day. Anyone who tells you otherwise is selling something.

This page explains, without a single technical word, what this tool does instead, how often it is wrong — with the measured number, which almost nobody publishes — and what you can do with any of it on Monday morning.

So what does it do?

Three separate things. What matters is not the three: it is that two of them are facts and one is an opinion, and confusing those is where the injustices come from.

An opinion about prose

The score. It names words and turns of phrase that machine writing leans on, and shows you each one underlined in the text. It is a judgement about how the writing sounds. It can be argued with, and sometimes it is wrong.

A fact

Characters nobody types. Tools that "humanise" text leave behind letters from another alphabet that look identical. It gives you the codepoint, the line and the column. Either it is there or it is not.

A fact

A bibliography against itself. A source cited in the text and missing from its own list; the same identifier on two different papers; a year that has not happened. No internet needed: the document contradicts itself.

A percentage can be argued about for half an hour in a meeting. The same identifier on two different references is settled with one question: "can you send me the paper?" That is why the tool shows the two apart, and never adds the second to the first.

How often is it wrong?

This is the question no detector answers, and the only one that matters when there is a person on the other side of it.

Here is how it was measured: 296 texts written before 2022 — published articles, Wikipedia revisions and classroom essays. Nobody judged them human; their dates did. Then they went through the tool and we counted how many it would have accused.

296 HUMAN TEXTS 2 flagged the two filled in, out of two hundred and ninety-six WHAT CAN HONESTLY BE CLAIMED 0.7% 0% 2.4% 10%
Two of the 296 were flagged: 0.7%. But two out of 296 is not exactly the rate either: with a sample that size, the most anyone can honestly claim is that it is wrong less than 2.4% of the time. That amber band is the difference between a measurement and a promise. It is not zero, and that is the point: until September this page showed a zero, because the corpus did not yet hold the people most likely to be accused falsely.

And it is not the same number for everybody

Handing the pooled figure to everyone would be convenient and false. The corpus holds 271 English texts and twenty-five Spanish ones, so what can be claimed about each language differs — and in the tool, each reader is shown their own, never the average.

English 271 texts 2.7% Spanish 25 texts 13.3% 0% 5% 10% 15%
Fewer Spanish texts means less certainty, not more mistakes. The bar is longer because we know less, and saying so is more use than hiding it behind an average.

What it does not tell you

How much machine writing it misses. That is not measured, deliberately: measuring it would mean collecting generated text, and any such collection represents whichever models were fashionable that month. A tool that flags nothing has a perfect error rate.

It does include student work, and that is the part that matters most to you. Of the 296 texts, 206 are classroom essays written by adults learning English. That is precisely the group detectors are accused of getting wrong most often, and here it is the group that fares worst: it sits highest of any group in every column. At a threshold of 25 out of 100, nine of them were flagged. That is why the boundary is 30 and not 25 — it moved once we knew who it was hurting.

What is still unmeasured is the student essay in Spanish: the twenty-five Spanish texts are Wikipedia revisions from a single source. There this tool knows less, and it says so rather than lending one language a figure it measured in another.

And one worth knowing before you look at any score: the signals that fire most often on human writing say almost nothing. Uniform sentence rhythm appears in 32.1% of the human texts we measured, and the word "just" in 34.8%. If the evidence against a piece of work leans mostly on those, it is worth less. The tool says so on every report, by name.

Why trust this one and not another?

Not because of the technology. Because of this: on 10 August 2026 this tool was accusing people falsely, and we published that.

A piece of work with a perfectly good bibliography, which merely listed some further reading without citing it, was announced in the report as "1 source contradiction", above the sentence "the document disagrees with itself". On the page a teacher prints and carries into a committee.

The bug was not the worst part. The code itself knew the difference: there is a field that separates a real contradiction from mere untidiness, with a comment explaining that people legitimately list further reading. The line printing the headline never asked it. It was published for five days and no user found it.

It is fixed, it is written up, and the history is public. Ask any paid detector when it last accused someone falsely. That answer, not a percentage, is what should decide who you trust.

What do I do on Monday?

None of this is worth anything until it becomes something you can say in a classroom or a meeting. That part is written, and there is no software in it:

Wording for your syllabus A one-page sheet for students A procedure for a committee

The student sheet carries the figure measured for their language, not the pooled one. If the work is in Spanish, the number that goes to the committee is 13.3%.

Try the tool Read the code

It runs entirely in your browser. The text is not uploaded anywhere — not even to us. This page requests nothing from any server either: no font, no image, no analytics.