typestar

word_freq.clj en Clojure

Cuenta palabras de un pasaje breve sobre un puerto e imprime las ocho más frecuentes con su porcentaje.

;; Word frequency report over one fixed passage.
(require '[clojure.string :as str])

(def passage
  "The harbor wakes early. Gulls circle the harbor wall while the
   ferry crews coil rope and the tide slides out. By noon the harbor
   is loud, and by dusk the harbor is quiet again.")

;; Splitting on runs of non-letters leaves an empty first piece when
;; the text opens with punctuation, so drop the blanks afterward.
(defn words [text]
  (->> (str/split (str/lower-case text) #"[^a-z']+")
       (remove str/blank?)))

;; frequencies builds a map of item -> count in a single pass.
(def counts (frequencies (words passage)))
(def total (count (words passage)))

;; sort-by can key on a vector: negate the count for descending
;; order, then fall back to the word itself to break ties.
(def ranked (sort-by (juxt (comp - val) key) counts))

(println (format "%-10s %5s %6s" "WORD" "COUNT" "SHARE"))
(println (apply str (repeat 23 "-")))
(doseq [[word n] (take 8 ranked)]
  (println (format "%-10s %5d %5.1f%%" word n (* 100.0 (/ n total)))))
(println (apply str (repeat 23 "-")))
(println (format "%-10s %5d" "WORDS" total))
(println (format "%-10s %5d" "DISTINCT" (count counts)))

Cómo funciona

  1. (str/split (str/lower-case text) #"[^a-z']+") parte el texto en minúsculas y descarta lo vacío.
  2. frequencies cuenta las palabras y (juxt (comp - val) key) las ordena por conteo y luego alfabético.
  3. (format "%-10s %5d %5.1f%%" word n (* 100.0 (/ n total))) imprime cada fila en columnas fijas bajo el encabezado.

Palabras clave y builtins usados aquí

El intento, en números

Líneas
29
Caracteres a escribir
1167
Tokens
171
Ritmo de tres estrellas
75 tpm

Al ritmo de tres estrellas de 75 tokens por minuto, este intento toma unos 137 segundos.

Escribe este fragmento

Paso 1 de 3 en Bis; paso 25 de 27 en Fundamentos del lenguaje.

← Anterior Siguiente →