word_freq.scala en Scala
Un reporte completo: normalizar un pasaje, contar sus palabras, ordenarlas e imprimirlas.
// A word-frequency report over an inline passage.
val passage: String =
"""Sailors watch the sky, because the sky warns the sailors.
|A red sky at night can delight a sailor; a red sky in the
|morning will warn the same sailor. Watch the sky, watch the
|wind, and the sea will rarely surprise you.""".stripMargin
// Lowercasing first means Sky and sky land in the same bucket.
// Splitting on runs of non-letters drops punctuation and newlines.
def words(text: String): List[String] =
text.toLowerCase.split("[^a-z]+").filter(_.nonEmpty).toList
// Two words carry no signal, so they are dropped from the report.
val stopWords: Set[String] = Set("the", "a", "and", "is", "in", "at", "you")
// groupBy(identity) buckets equal words; the size of a bucket is
// that word's count.
def tally(items: List[String]): Map[String, Int] =
items.groupBy(identity).view.mapValues(_.size).toMap
@main def wordFreq(): Unit =
val counted = tally(words(passage).filterNot(stopWords.contains))
// Sorting by the negated count puts the most frequent first, and
// the word itself breaks ties alphabetically.
val ranked = counted.toList.sortBy((word, count) => (-count, word))
val distinct = counted.size
val total = counted.values.sum
println(f"${"word"}%-12s${"count"}%6s${"share"}%8s")
println("-" * 26)
for (word, count) <- ranked.take(8) do
val share = 100.0 * count / total
println(f"$word%-12s$count%6d$share%7.1f%%")
println("-" * 26)
println(f"${"total"}%-12s$total%6d")
println(f"$distinct distinct words kept out of ${words(passage).size}")
Cómo funciona
text.toLowerCase.split("[^a-z]+")borra mayúsculas y puntuación en una sola pasada.items.groupBy(identity).view.mapValues(_.size).toMapvuelve conteos la lista de palabras.sortBy((word, count) => (-count, word))ordena por conteo y luego alfabético;$word%-12s$count%6dalinea columnas.
Palabras clave y builtins usados aquí
defdoforval
El intento, en números
- Líneas
- 37
- Caracteres a escribir
- 1538
- Tokens
- 262
- Ritmo de tres estrellas
- 75 tpm
Al ritmo de tres estrellas de 75 tokens por minuto, este intento toma unos 210 segundos.
Paso 1 de 3 en Bis; paso 25 de 27 en Fundamentos del lenguaje.