TopPodcast.com
Menu
  • Home
  • Top Charts
  • Top Networks
  • Top Apps
  • Top Independents
  • Top Podfluencers
  • Top Picks
    • Top Business Podcasts
    • Top True Crime Podcasts
    • Top Finance Podcasts
    • Top Comedy Podcasts
    • Top Music Podcasts
    • Top Womens Podcasts
    • Top Kids Podcasts
    • Top Sports Podcasts
    • Top News Podcasts
    • Top Tech Podcasts
    • Top Crypto Podcasts
    • Top Entrepreneurial Podcasts
    • Top Fantasy Sports Podcasts
    • Top Political Podcasts
    • Top Science Podcasts
    • Top Self Help Podcasts
    • Top Sports Betting Podcasts
    • Top Stocks Podcasts
  • Podcast News
  • About Us
  • Podcast Advertising
  • Contact
Not in our directory?
Add Show Here
Podcast Equipment
Center

toppodcastlogoOur TOPPODCAST Picks

  • Comedy
  • Crypto
  • Sports
  • News
  • Politics
  • True Crime
  • Business
  • Finance

Follow Us

toppodcastlogoStay Connected

    View Top 200 Chart
    Back to Rankings Page
    Technology

    Functional Design in Clojure

    Each week, we discuss a software design problem and how we might solve it using functional principles and the Clojure programming language.

    Advertise

    Copyright: © 2018-2019, Christoph Neumann and Nate Jones

    • Apple Podcasts
    • Google Play
    • Spotify

    Latest Episodes:
    Ep 039: Why Use Clojure Over Another Functional Language? Jul 26, 2019
    Show notes

    Each week, we answer a different question about Clojure and functional programming.

    If you have a question you'd like us to discuss, tweet @clojuredesign, send an email to feedback@clojuredesign.club, or join the #clojuredesign-podcast channel on the Clojurians Slack.

    This week, the question is: "Why use Clojure over another functional language?". We examine the different categories of functional programming languages and distill out what differentiates Clojure and why we prefer it.

    Selected quotes:

    • "Running just one function when developing is not only allowed in Clojure, it's encouraged and celebrated."
    • "You don't have to make the whole world (application) agree. You can work on just a part of it and then bring it back into the rest of the world when you want it to agree."
    • "I would like some XML in my cake."
    • "Oh, you were a hipster Scala user."
    • "When I pull in code off clojars, it's going to use the Clojure way, because there is a Clojure way."
    • "If you can make all your abstractions with a simpler set of semantics, wouldn't that be better than a broader set?"
    • "Multi-paradigm languages are inherently more complex. You really end up in the 'good parts' kind of problem. Scala, The Good Parts. Javascript, The Good Parts."
    • "Code is about communicating with two things. The computer and the other developers. The computer can handle esoteric language features, but other developers will have a harder time with them."

    Related episodes:

    • 002: Tic-Tac-Toe, State in a Row
    • 014: Fiddle with the REPL

    Ep 038: How Do I Convince My Coworkers to Use Clojure? Jul 19, 2019
    Show notes

    Each week, we answer a different question about Clojure and functional programming.

    If you have a question you'd like us to discuss, tweet @clojuredesign, send an email to feedback@clojuredesign.club, or join the #clojuredesign-podcast channel on the Clojurians Slack.

    This week, the question is: "How do I convince my coworkers to use Clojure?". We recall our own experiences evangelizing Clojure and give practical advice from the trenches.

    Selected quotes:

    • "Don't assume someone is going to want to jump on the grenade."
    • "There are actually two steps: 1. Convince people that Clojure is a good thing. 2. Get people to actually use it."
    • "Lambda as in AWS Lambda, not as in Clojure lambda. They stole our cool word!"
    • "By the way, that solution came from Clojure."
    • "Get people thinking in the functional direction."
    • "There are only two types of tools: those that you hate and those you don't use."
    • "When people get stuck, they think it's the language's fault, because it has chosen to categorically eliminate solutions that they, as a developer, have relied on for years."
    • "The path to Clojure adoption is primarily social, not technical."

    Related episodes:

    • 024: You Are Here, But Why?

    Ep 037: What Advice Would You Give to Someone Getting Started With Clojure? Jul 12, 2019
    Show notes

    Each week, we answer a different question about Clojure and functional programming.

    If you have a question you'd like us to discuss, tweet @clojuredesign, send an email to feedback@clojuredesign.club, or join the #clojuredesign-podcast channel on the Clojurians Slack.

    This week, the question is: "What advice would you give to someone getting started with Clojure?". We trade off giving practical tips for intrepid learners while we reminisce about our own paths into Clojure.

    Selected quotes:

    • "Reading other people's code will change how you view the problem."
    • "You're going to get stuck, but that's ok.
    • "REPL driven development is so fast, you'll need to take a break from the whiplash."
    • "Clojure doesn't give you breaks, like other programming languages do with their compiles and restarts."
    • "The best code to use is code that was written by somebody else, because it's probably bug free."
    • "The problem isn't the code that I'm writing, the problem is with my approach."
    • "Try to solve a programming problem that you care about."

    Related episodes:

    • 012: Embrace the REPL
    • 013: Connect the REPL
    • 014: Fiddle with the REPL

    Ep 036: Why Do You Recommend Clojure? Jul 05, 2019
    Show notes

    It's summertime, and that means it's time for something new. Each week, we will answer a different question about Clojure and functional programming.

    If you have a question you'd like us to discuss, tweet @clojuredesign or send an email to feedback@clojuredesign.club.

    This week, we're starting off with "Why do you recommend Clojure?". We take turns sharing our favorite reasons, and we can't help but have fun riffing on how enjoyable Clojure is to use. Come along for the ride.

    Selected quotes:

    • "Question everything else. Clojure is the answer."
    • "Once you discover structural editing, you'll never go back."
    • "Data is inert. It can't act."
    • "When dealing with mutable data and side effects, it's like trying to add a domino in the middle and make sure to not break the chain."
    • "Most of the time, running with scissors is considered to be a bad thing."
    • "Somebody in Java has done all that plumbing that you don't really want to do."

    Ep 035: Lifted Learnings Jun 28, 2019
    Show notes

    Christoph and Nate lift concepts from the raw log-parsing series.

    • Reflecting on the lessons learned in the log series.
    • (01:15) Concept 1: We found Clojure to be useful for devops.
      • Everything is a web application these days,
      • "The only UIs in Devops are dashboards."
      • For most of the series, our UI was our connected editor.
      • We grabbed a chunk of the log file and were fiddling with the data in short order.
      • We talk about connected editors in our REPL series, starting with Episode 12.
      • Being able to iteratively work on the log parsing functions in our editor was key to exploring the data in the log files.
    • (04:04) Concept 2: Taking a lazy approach is essential when working with a large data set.
      • Lazily going through a sequence is reminiscent of database cursors. You are at some point in a stream of data.
      • We ran into some initial downsides.
      • When using with-open, fully lazy processing results in an I/O error, because the file has been closed already.
      • Shouldn't be too eager too early, because then the entire dataset will reside in memory.
      • Two kinds of functions: lazy and eager.
        • Lazy functions only take from a sequence as they need more values.
        • Eager functions consume the whole sequence before returning.
      • Ensure that only the last function in the processing chain is eager.
      • "It only takes one eager to get everybody unlazy."
    • (08:38) Concept 3: Clojure helps you make your own lazy sequences using lazy-seq.
      • Clojure has a deep library of functions for making and processing lazy sequences.
      • We were able to make our own lazy sequences that could then be used with those functions.
      • Wrap the body in lazy-seq and return either nil (to indicate the end) or a sequence created by calling cons on a real value and a recursive call to itself.
    • (12:41) Concept 4: We work with information at different levels, and that forms an information hierarchy.
      • The data goes from bits to characters to lines, and then we get involved.
      • We move from lines on up to more meaningful entities. Parsed lines are maps that have richer information, and then errors are richer still.
      • Our parsers take a sequence and emit a new sequence that is at a higher level of information.
      • We first explored this concept in the Time series.
      • The transformations from one level to the next are all pure.
    • (14:53) Concept 5: Sometimes you have to go down before you can go up again another way.
      • We pre-abstracted a little bit, and only accepted lines that had all of the data we were looking for (time, log level, etc.).
      • Exceptions broke that abstraction, so we reworked our "parsed line" map to make the missing keys optional.
    • (15:54) Concept 6: Maps are flexible bags of dimensions. They are a set of attributes rather than a series of rigid slots that must be filled.
      • Functions only need to look at the parts of the map that they need.
      • Every time we amplify the data, we add a new set of dimensions.
      • Thanks to namespacing, all of these dimensions coexist peacefully.
      • Multiple levels of dimensions give you more to filter/map/reduce on.
      • Just because you distill, doesn't mean you want to lose essence.
    • (21:09) Concept 7: Operating within a level of information is a different concern than lifting up to a higher level of information.
      • Within a level, functions aid in filtering and aggregating.
      • Between levels, functions recognize patterns and groupings to produce higher levels of information.
      • Make the purpose of the function clear in how you name it.
      • Separate functions that "lift" the data from functions that operate at the same level of information.
      • When exploring data, you don't know where it will lead, so start by moving the data up a level in small steps.

    Related episodes:

    • 012: Embrace the REPL
    • 015: Finding the Time
    • 028: Fail Donut
    • 029: Problem Unknown: Log Lines
    • 030: Lazy Does It
    • 031: Eager Abstraction
    • 032: Call Me Lazy
    • 033: Cake or Ice Cream? Yes!
    • 034: Break the Mold

    Clojure in this episode:

    • lazy-seq, cons
    • with-open

    Ep 034: Break the Mold Jun 21, 2019
    Show notes

    Christoph finds exceptional log lines and takes a more literal approach.

    • Previously, we upgraded our log parsing to handle finding two errors instead of just one.
    • "It's amazing what you don't find when you're not looking."
    • We ended up with a set of functions that can parse out multiple errors.
    • The result is a nice mixed sequence of errors that we can aggregate and pivot with Clojure core.
    • "By the power of Clojure core, I will map and reduce!"
    • (02:43) New Problem: there are exceptions in the log, and they span multiple lines!
    • The exception continuation lines don't parse.
    • They don't have a date, a log level, or anything else that regular log lines have.
    • How do we collect up all those lines? They are all part of one logical error.
    • Our error parsers don't get a chance to see these lines, because the general line parser threw them away.
    • (06:27) Solution 1: Don't pre-parse the lines, but instead have each error parsing function do the general parse.
    • Each parsing function would receive the raw, unparsed, line as a string.
    • Each function would have to run parse-line on the inputs before doing their specific parsing.
    • Now each function in our inventory needs to re-do the general parsing of the line.
    • This ends up being much more inefficient.
    • (10:13) Solution 2: Relax the general parsing.
    • What is the general parser to do with a line that doesn't match the regexp?
    • "How about if the general parse does the best job it can, and whatever it can't do, it doesn't do?"
    • One of the keys in the data returned from parse-line is :raw/line, which is the entire line, and that's always there.
    • So, if the regexp fails, we can always return at least that.
    • The map doesn't have the keys that are not found.
    • Now, parse-line will always return a map for each line.
    • We've uplifted the data, at least slightly, from a string to a map.
    • We can know two things
      • Every map will have :raw/line
      • If a map doesn't have the other keys (like :log/date), that means it didn't parse, and is probably a continuation of a previous line.
    • This is us expanding our program's view of reality to more closely match actual reality.
    • We left out part of reality, and got stuck because of it.
    • Being more literal to the data source gives us more flexibility.
    • Clojure's dynamic nature shines in this situation. We don't have to have all the keys all the time.
    • "A map is a bucket of dimensions."
    • Each function that operates on the data will inspect the data and see if it has all the keys it needs. It doesn't matter if there are other keys.
    • Namespaced keys really help in this situation, each new level of parsing can add keys without overwriting keys from other dimensions.
    • How does this impact our parsing functions?
    • They are unchanged, because they grab the message out using some->>, which shortcuts on nil.
    • (19:08) Our exception error parsing function detect the start in a number of ways.
      • The next line is a bare line?
      • The first line ends in an opening curly brace?
    • Then, to find the end of the exception, it can use take-while to find all the lines that are bare.
    • To assist in finding bare lines, we can introduce a predicate function instead of embedding that logic.
    • When all lines are found, it then combines them into a new :log/message key.
    • When updating the error map in our parsing functions, we are careful to only operate on the data we know is there, so that other keys are not impacted.

    Related episodes:

    • 028: Fail Donut
    • 029: Problem Unknown: Log Lines
    • 030: Lazy Does It
    • 031: Eager Abstraction
    • 032: Call Me Lazy
    • 033: Cake or Ice Cream? Yes!

    Clojure in this episode:

    • some->>, take-while

    Code sample from this episode:

    (ns devops.week-06
      (:require
        [clojure.string :as string]
        [devops.week-02 :refer [process-log]]
        [devops.week-05 :refer [parse-357-error parse-sprinkle]]
        ))
    
    
    (def general-re #"(\d\d\d\d-\d\d-\d\d)\s+(\d\d:\d\d:\d\d)\s+\|\s+(\S+)\s+\|\s+(\S+)\s+\|\s+(\S+)\s+\|\s(.*)")
    
    (defn parse-line
      [line]
      (if-let [[whole dt tm thread-name level ns message] (re-matches general-re line)]
        {:raw/line whole
         :log/date dt
         :log/time tm
         :log/thread thread-name
         :log/level level
         :log/namespace ns
         :log/message message}
        {:raw/line line
         :log/message line}))
    
    (defn bare-line?
      [line]
      (nil? (:log/date line)))
    
    (defn parse-exception-info
      [lines]
      (let [first-line (first lines)
            [_whole classname] (some->> first-line :log/message (re-matches #"([^ ]+) #error \{"))]
        (when classname
          (let [error-lines (cons first-line (take-while bare-line? (rest lines)))
                error-str (string/join "\n" (map :log/message error-lines))]
            (merge first-line
                   {:kind :error
                    :error/class classname
                    :log/message error-str})))))
    
    (defn parse-next
      [lines]
      (or (parse-357-error lines)
          (parse-sprinkle lines)
          (parse-exception-info lines)))
    
    (defn parse-all
      [lines]
      (lazy-seq
        (when (seq lines)
          (if-some [found (parse-next lines)]
            (cons found (parse-all (rest lines)))
            (parse-all (rest lines))))))
    
    
    (comment
      (process-log "sample.log" #(->> % (map parse-line) parse-all doall))
      )

    Log file sample:

    2019-05-14 16:48:57 | process-Poster | INFO  | com.donutgram.poster | transaction failed while updating user joe: code 357
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | failed to add sprinkle to donut 50493
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | sprinkle fail reason: unknown state
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | Poster #error {
     :cause "Failed to lock the synchronizer"
     :data {}
     :via
     [{:type clojure.lang.ExceptionInfo
       :message "Failed to lock the synchronizer"
       :data {}
       :at [process.poster$eval50560 invokeStatic "poster.clj" 40]}]
     :trace
     [[process.poster$eval50560 invokeStatic "poster.clj" 40]
      [clojure.lang.AFn run "AFn.java" 22]
      [java.lang.Thread run "Thread.java" 748]]}

    Ep 033: Cake or Ice Cream? Yes! Jun 14, 2019
    Show notes

    Nate needs to parse two different errors and takes some time to compose himself.

    • Previously, we were able to parse out errors and give the parsing function the ability to search as far into the future as necessary.
    • We did this by having the function take a sequence and return a sequence, managed by lazy-seq.
    • (01:30) New Problem: We need to correlate two different kinds of errors.
    • The developers looked at our list of sprinkle errors and they think that they're caused by the 357 errors.
    • They have requested that we look at the entire log and generate a report of 357 and sprinkle errors, so we can tell if they're correlated.
    • "When someone says, do I want cake or ice cream, the right answer is: yes, I want both!"
    • Before, we were only parsing out a single type of error and summarizing it, but now we need to parse out both types of errors.
    • If we try to parse both kinds of errors with the same function, we will quickly get ourselves into nested ifs or maybe an infinite cond. Perhaps a complex state machine with backtracking?
    • (05:55) Realization: Each error stands alone. Once you detect the beginning of a sprinkle error, you won't need to look for a 357 error.
    • You can take each one in turn.
    • (06:30) Solution step 1: What if we had two functions, one for each type of error.
    • Each of these functions would take the entire sequence and tell us if there was an error at the beginning.
    • Previously, our function both recognized errors and handled the sequence generation. If we pull those apart, we can add parsing for more errors easily.
    • Each error parsing function would return nil if no error was found at the head of the sequence.
    • (08:46) Solution step 2: Create a function that uses the two detectors to find out what error is at the head of the sequence.
    • It takes the sequence, and wraps consecutive calls in an or block.
    • The or block will try each one in turn until one matches and then that is the result.
    • Each error's parsing is in its own function, and the combining function serves as an inventory.
    • (11:35) Solution step 3: Create a lazy sequence that wraps calls to the combined detector function.
    • Last week's code has parsing and lazy in one function.
    • Now that we've pulled the parsing out, we can use the remaining structure to create our lazy sequence.
    • The combined detector function is parse-next, and the function that manages the lazy sequence is parse-all.
    • "Now we've fulfilled our obligation to have bike-shedding on naming. Next up, cache consistency. And finally, off-by-one errors."
    • The top of parse-all has a call to lazy-seq.
    • It will use the result of calling parse-next on the sequence.
      • If it gets something, it will use cons to add that value to the beginning of a recursive call to itself.
      • If it gets nil, it will recursively call itself with the rest of the sequence, thus advancing the parsing forward one step.
    • It's not a ton of boilerplate, but it is nice to put all the mechanics of the sequence creation into a function by itself.
    • Now we have a heterogeneous sequence of errors, and we can transform it into any report that is useful.
    • Each parsing function doesn't need to worry about advancing down the sequence, that is handled by the higher parse-all function.
    • Since we have a new lazy sequence, we can take it and make recognizers that take it and generate an even higher level of sequence.
    • We ruminate more on higher level data in Episode 020.

    Related episodes:

    • 020: Data Dessert
    • 028: Fail Donut
    • 029: Problem Unknown: Log Lines
    • 030: Lazy Does It
    • 031: Eager Abstraction
    • 032: Call Me Lazy

    Clojure in this episode:

    • seq, cons, rest
    • lazy-seq
    • or, cond

    Code sample from this episode:

    (ns devops.week-05
      (:require
        [devops.week-01 :refer [parse-line]]
        [devops.week-02 :refer [process-log]]
        [devops.week-03 :refer [sprinkle-errors-by-type]]
        ))
    
    (defn parse-sprinkle
      [lines]
      (let [[first-line second-line] lines
            [_whole donut-id] (some->> first-line :log/message (re-matches #"failed to add sprinkle to donut (\d+)"))
            [_whole error] (some->> second-line :log/message (re-matches #"sprinkle fail reason: (.*)"))]
        (when (and donut-id error)
          (merge first-line
                 {:kind :sprinkle
                  :sprinkle/donut-id donut-id
                  :sprinkle/error error}))))
    
    (defn parse-357-error
      [lines]
      (let [[first-line] lines
            [_whole user] (some->> first-line :log/message (re-matches #"transaction failed while updating user ([^:]+): code 357"))]
        (when user
          (merge first-line
                 {:kind :code-357
                  :code-357/user user}))))
    
    (defn parse-next
      [lines]
      (or (parse-357-error lines)
          (parse-sprinkle lines)))
    
    (defn parse-all
      [lines]
      (lazy-seq
        (when (seq lines)
          (if-some [found (parse-next lines)]
            (cons found (parse-all (rest lines)))
            (parse-all (rest lines))))))
    
    (defn kind?
      ([kind]
       #(kind? % kind))
      ([line kind]
       (= kind (:kind line))))
    
    
    (comment
      (process-log "sample.log" #(->> % (map parse-line) parse-all (map :kind) doall))
      (process-log "sample.log" #(->> % (map parse-line) parse-all (filter (kind? :sprinkle)) sprinkle-errors-by-type))
      )

    Log file sample:

    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | transaction failed while updating user joe: code 357
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | failed to add sprinkle to donut 23948
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | sprinkle fail reason: should never happen
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | failed to add sprinkle to donut 94238
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | sprinkle fail reason: timeout exceeded threshold
    2019-05-14 16:48:56 | process-Poster | INFO  | com.donutgram.poster | transaction failed while updating user sally: code 357
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | failed to add sprinkle to donut 24839
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | sprinkle fail reason: too many requests
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | failed to add sprinkle to donut 19238
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | sprinkle fail reason: should never happen
    2019-05-14 16:48:57 | process-Poster | INFO  | com.donutgram.poster | transaction failed while updating user joe: code 357
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | failed to add sprinkle to donut 50493
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | sprinkle fail reason: unknown state

    Ep 032: Call Me Lazy Jun 07, 2019
    Show notes

    Christoph finds map doesn't let him be lazy enough.

    • Last week, we were dealing with multi-line sprinkle errors.
    • We were able to get more context using partition.
    • (01:33) Problem: the component lines had to be adjacent.
    • Solution last week was to create larger partitions to hopefully get the rest of the error.
    • This became a magic number problem, guessing how far we had to look ahead.
    • "If there's anything I've learned in my career, telling the future is one of the hardest things to do.'
    • What number should be big enough? 100? 1000?
    • (04:00) The other problem is that the function is handed a pre-selected set of lines.
    • The decision about how many lines is appropriate is made outside the function.
    • Wouldn't it be nice if the function had control over how far to look ahead.
    • "The function can't function."
    • "Functions are all we've got in functional programming. Well, that and lists."
    • It would be great if the function itself could take a sequence and look as far as it needs to.
    • How about handing the function the entire lazy sequence?
    • (05:52) Problem: Handing in the entire sequence means we can't use map to convert lines into sprinkle errors anymore.
    • We can write a function that gives us just one sprinkle error from the sequence, but we want to convert the sequence into a sequence of all the sprinkle errors.
    • We're going from something that operates on a subset of the sequence to something that operates on the entire sequence, which is too much control.
    • We need a way for it to look ahead but still
    • It's no longer just working on a chunk of the sequence, but on the unbounded sequence itself.
    • We need to elevate it to the same power as other sequence operators, like map and filter.
    • We don't, however, want the function to eagerly find all sprinkle errors in the sequence. It needs to be lazy.
    • (09:26) Solution 1: How can we just get one sprinkle error out?
      1. If the first line isn't the error start, recur with the tail until found.
      2. Do a take-while to find the second half of the error.
      3. When both found, return the value.
    • We need to terminate the search if we hit the end of the sequence, so we only continue if (seq lines) is not nil.
    • "There's no sense in looking in an empty bucket."
    • But we don't want just one, we want the entire sequence.
    • It would be really nice to return the value when we find it and then wait to find the next one until it is requested.
    • Conceptually, we could tell the calling function an index of where to start looking for the next error.
    • (13:44) Solution: In Clojure, we keep our place using the lazy-seq function.
    • lazy-seq is a sequence, but it hasn't been realized yet.
    • It's like being able to hand back a value and a function to call for the next value.
    • When you find a value, you can cons it onto the head of an invocation of lazy-seq to make a new sequence.
    • Step 1. Wrap your entire function body in lazy-seq.
    • This is similar to using delay, because it wraps the code in something that will only be evaluated when it is first accessed.
    • Step 2. Ensure that the body obeys the contract. It must return either:
      • nil, which indicates that the sequence is complete.
      • a sequence, usually constructed by calling cons on a value and a call to lazy-seq.
    • Top of the body is a call to (when (seq lines) ..., to ensure that the sequence terminates when there is no data left.
    • Since the top of our function is lazy-seq, we can cons the found value onto a recursive call to the function.
    • In the recursive call, we must pass the next section of the sequence, so that when evaluated it will pick up at the right place.
    • If we don't find the start of the error, we recurse with the rest of the sequence to try parsing from there.
    • This function will go through the sequence eagerly until it finds something.
    • Instead of operating on single elements in the sequence, we can take a sequence and produce a sequence, powered by lazy-seq.
    • With this capability, you can build a higher level sequence that consumes this sequence and produces a new summary, all done lazily.

    Related episodes:

    • 028: Fail Donut
    • 029: Problem Unknown: Log Lines
    • 030: Lazy Does It
    • 031: Eager Abstraction

    Clojure in this episode:

    • partition
    • seq, cons, rest
    • lazy-seq, delay
    • map, filter, take-while
    • recur

    Code sample from this episode:

    (ns devops.week-04
      (:require
        [devops.week-01 :refer [parse-line]]
        [devops.week-02 :refer [process-log]]
        [devops.week-03 :refer [sprinkle-errors-by-type]]
        ))
    
    (defn sprinkle-error-seq
      [lines]
      (lazy-seq
        (when (seq lines)
          (let [[first-line second-line & tail] lines
                [_whole donut-id] (some->> first-line :log/message (re-matches #"failed to add sprinkle to donut (\d+)"))
                [_whole error] (some->> second-line :log/message (re-matches #"sprinkle fail reason: (.*)"))]
            (if (and donut-id error)
              (cons (merge first-line
                           {:kind :sprinkle
                            :sprinkle/donut-id donut-id
                            :sprinkle/error error})
                    (sprinkle-error-seq tail))
              (sprinkle-error-seq (next lines)))))))
    
    
    (comment
      (process-log "sample.log" #(->> % (map parse-line) sprinkle-error-seq doall))
      (process-log "sample.log" #(->> % (map parse-line) sprinkle-error-seq sprinkle-errors-by-type))
      )

    Ep 031: Eager Abstraction May 31, 2019
    Show notes

    Nate finds that trouble comes in pairs.

    • Last week, we were able to make our parsing lazy.
    • Each step was lazy, and we side-stepped the landmine of waiting to perform computation till after the file is closed.
    • (02:00) There's a new error: Sprinkles are failing!
    • Sprinkles are the "likes" of DonutGram, and are the metric by which all donuts are judged.
    • Log lines are showing up, like failed to add sprinkle to donut 23948.
    • Let's write another regexp and handler. This is easy!
    • But wait, we notice that there is another log line right after: sprinkle fail reason: db timeout exceeded.
    • We want to capture that reason as part of the error, but it's on a different line.
    • Our current code (from Episode 029), has a nifty abstraction with a paired regexp and handler function. If the regexp matches, then the matches are passed to the handler.
    • The underlying assumption is that each line stands alone.
    • Now we have two lines, and that breaks our pre-abstraction.
    • Our function needs more context, give our function two lines instead of just one.
    • Our parse-details function goes out the window, what can we do to get more context for each parsing?
    • (07:05) Let's go to the Clojure library and check out the partition function.
    • Used to break up a list into chunks. Eg. Chunking arg lists into key/value pairs.
    • The step argument varies how far you reach into the collection for the next chunk.
    • For instance, (partition 2 1 (range 1 7)) yields ((1 2) (2 3) (3 4) (4 5) (5 6)). Each element is paired with it's following element.
    • Doing this with our log lines means we have a second line in the group that we hand to our parsing function.
    • The parsing function can look in the first line for the error, and then if it finds it, look in the second line for the reason. If both are found, we return a sprinkle error record.
    • If either regex fails, nil is returned, so our list becomes a sequence of sprinkle errors or nils.
    • Then we can filter out the nils and summarize the sprinkle errors.
    • "Some of these errors hurt more than others."
    • The sequence is not fully realized, because partition is lazy. It constructs chunks on demand.
    • "This is the beauty of Clojure core, all this stuff is lazy. Except, of course, for the eager parts."
    • (15:45) Problem: when looking through log lines, there are often other log lines in between.
    • What happens if our sprinkle fail reason is not on the line immediately after the sprinkle error?
    • We might need to look ahead a little further ahead than the next line.
    • One solution is to increase the partition size when chunking the lines. How about (partition 10 1 lines) or (partition 100 1 lines)?
    • The first line is the indication that there's an error, and then look through the rest of the lines for the reason.
    • We could log out "should never happen" if we don't find the error in the next 10 lines.
    • It becomes a magic number problem, how far do we need to look ahead?
    • At least these large partitions use structural sharing to re-use the record memory. We only pay extra for the containing collections.
    • "Is there a partition-infinity?"
    • "It feels like something we can be lazy about, and push off to our next episode."
    • "It's like when a show jumps the shark and you know it will come to an end, but you're just not sure which season that's going to be in."
    • What if parsing the first line determines how many further lines we need to consume?
    • "Fun is a loaded term."
    • "It always puts you in an interesting situation when your pre-optimization doesn't quite hold up a month down the road."
    • "I had the perfect abstraction for the wrong shape."
    • "The world changed around me." "Clearly it's not the fault of my code."

    Related episodes:

    • 028: Fail Donut
    • 029: Problem Unknown: Log Lines
    • 030: Lazy Does It

    Clojure in this episode:

    • partition
    • ->>
    • map, filter
    • group-by, frequencies

    Code sample from this episode:

    (ns devops.week-03
      (:require
        [devops.week-01 :refer [parse-line]]
        [devops.week-02 :refer [process-log]]
        ))
    
    (defn parse-sprinkle-error
      [line-pairs]
      (let [[first-line second-line] line-pairs
            [_whole donut-id] (some->> first-line :log/message (re-matches #"failed to add sprinkle to donut (\d+)"))
            [_whole error] (some->> second-line :log/message (re-matches #"sprinkle fail reason: (.*)"))]
        (when (and donut-id error)
          (merge first-line
                 {:kind :sprinkle
                  :sprinkle/donut-id donut-id
                  :sprinkle/error error}))))
    
    (defn sprinkle-errors
      [lines]
      (->> lines
           (partition 2 1)
           (map parse-sprinkle-error)
           (remove nil?)))
    
    (defn sprinkle-errors-by-type
      [errors]
      (->> errors
           (map :sprinkle/error)
           (frequencies)))
    
    
    (comment
      (process-log "sample.log" #(->> % (map parse-line) sprinkle-errors sprinkle-errors-by-type))
      )

    Log file sample:

    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | failed to add sprinkle to donut 23948
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | sprinkle fail reason: should never happen
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | failed to add sprinkle to donut 94238
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | sprinkle fail reason: timeout exceeded threshold
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | failed to add sprinkle to donut 24839
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | sprinkle fail reason: too many requests
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | failed to add sprinkle to donut 19238
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | sprinkle fail reason: should never happen
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | failed to add sprinkle to donut 50493
    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | sprinkle fail reason: unknown state

    Ep 030: Lazy Does It May 24, 2019
    Show notes

    Christoph's eagerness to analyze the big production logs shows him the value of being lazy instead.

    • Last time: going through the log file looking for the mysterious 'code 357'.
    • "The error message that just made sense to the person who wrote it. At the time written. For a few days"
    • Back and forth with the dev team, but our devops sense was tingling.
    • Took a sample, fired up a REPL,
    • Ended up with a list of tuples:
      • First element: regexp to match
      • Second element: handler to transform matches into data
    • (02:00) It's running slower and slower, the bigger the log file we analyze.
    • "This is a small 4-person tech company." "Where we can use new technologies in the same decade that they were created?" "Yes!"
    • Problem: No one turned on log rotation! The log file is 7G and our application crashes.
    • "I think we should get lazy."
    • "Work harder by getting lazier."
    • "Haskell was lazy before it was cool!"
    • Each line contains all the information we need, so we can process them one at a time.
    • (4:30) Eager and lazy is like the difference between push and pull.
    • Giving someone a big bag of work to do, or having them grab more work as they finish.
    • The thing doing the work needs a little more to work on, so it pulls it in.
    • Clojure helps with this. It gives us an abstraction so we don't have to see the I/O happening when we do our processing.
    • When your map of a lazy sequence needs more data, it gets it on demand.
    • Clojure core is built to support this lazy style of processing.
    • File I/O in Java is lazy in the same way. It reads data into a buffer and when that buffer is used up, more is read.
    • Lazy processing is like a bucket brigade. At the head, you pour out your bucket and the person next to you notices the empty bucket and fills it up. Then this is repeated down the line as each bucket demands to be filled.
    • (07:55) Let's make our code lazy.
    • Current lines function slurps in the file and splits on newline.
    • Idea: Convert it to open the file and return a lazy sequence using line-seq.
    • The return value can be threaded through the rest of our pipeline.
    • Each step of our pipeline is lazy, and the start is lazy, so the whole process should be lazy.
    • "It's lazy all the way."
    • Problem: We run it, and BOOM, we get an I/O error.
    • What happened? We were using the with-open macro, which closes the file handle after the body is complete.
    • Since the I/O is delayed till we consume the sequence, when we start the file is already closed.
    • "The ability to pull more I/O out of the file has been terminated."
    • "Nothing at all, which is a lot less useful than something."
    • (12:29) Rather than having a lines function, why don't we just bring that code into the summary function?
    • Entire process is wrapped in a with-open so that all steps including summary complete before the file is closed.
    • Takes a filename and returns an incident count by user.
    • It does all that we want, but it's too chunky. We're doing too much in the function.
    • We usually want to move I/O to the edges and this commingles it with our logic.
    • "We just invited I/O to come move into the middle of our living room."
    • I/O and summary are in the same function, so to make a new summary, we have to duplicate everything.
    • We could split out the guts, extract the general and detailed parsing into a separate function. For reuse.
    • This means you are only duplicating the with-open and line-seq for each new summary.
    • (16:29) How can we stop duplicating the with-open? To separate that idiom into just one place.
    • If you can't return the line-seq, is there a way we can hand in the logic we need all at once?
    • Idea: Make a function that does the with-open and takes a function.
    • "Let's make it higher order."
    • We hand in the "work to do" as a function.
    • "What should we call it? How about process. That's nice and generic."
    • We turn the problem of abstraction into the problem of writing a function that takes a line sequence and produces a result.
    • Any functions that take and produce a sequence, including those that we wrote, can be used to write this function.
    • Clojure gives us several ways of composing functions together, we'll use ->> (the thread-last macro) in this case.
    • As we improve the vocabulary for working with lines, our ability to express the function gains power and conciseness.
    • Design tension: if there is something that needs to be done for every summary, it can be pushed into the process function.
    • The downside to that is that we sign up for that for every summary, and that might not be appropriate for the ones we haven't written yet.
    • We opt for making it easier to express and compose the function passed in.
    • We can still make a function that takes a filename and returns a summary, but the way we construct that function is through composition of transforms.
    • We can pre-bake a few transforms into shorter names if we want to use them repeatedly.
    • (23:20) We will still run into the I/O error problem if we're not careful.
    • The function that we pass to process needs to have an eager evaluation at the end.
    • If all we do is transform with lazy functions, the I/O won't start before the list is returned.
    • group-by or frequencies will suffice, but if you don't have one of those, reach for doall.
    • "You gotta give it a good swift kick with doall."
    • Style point: doall at the beginning or at the end of the thread? We like it at the end.
    • (26:07) We have everything we need.
      • Lazy so we don't pull in the entire file.
      • I/O sits in one function.
      • We have control over when we're eager.

    Message Queue discussion:

    • (26:38) Long-time listener Dave sent us a code sample!
    • An alternative implementation for parse-details that doesn't use macros.
    • Top level is an or macro.
    • Inside the or, each regex is matched in a when-let, the body of which uses the matches to construct the detailed data.
    • If the regex fails to match, nil is returned and the or will move on to the next block.
    • We tend to think of or as only for booleans, but it works well for controlling program flow as well.
    • The code is very clean and concise. And it only uses Clojure core.
    • "Without dipping into macro-land... Not that there's anything wrong with that."

    Related episodes:

    • 028: Fail Donut
    • 029: Problem Unknown: Log Lines

    Clojure in this episode:

    • slurp, with-open, line-seq
    • ->>
    • or, when-let
    • map, filter
    • group-by, frequencies
    • doall
    • clojure.string/split-lines

    Code sample from this episode:

    (ns devops.week-02
      (:require
        [clojure.java.io :as io]
        [devops.week-01 :refer [parse-line parse-details]]
        ))
    
    
    ; Parsing and summarizing
    
    (defn parse-log
      [raw-lines]
      (->> raw-lines
           (map parse-line)
           (filter some?)
           (map parse-details)))
    
    (defn code-357-by-user
      [lines]
      (->> lines
           (filter #(= :code-357 (:kind %)))
           (map :code-357/user)
           (frequencies)))
    
    
    ; Failed Attempt: returning from with-open
    
    (defn lines
      [filename]
      (with-open [in (io/reader filename)]
        (line-seq in)))
    
    (defn count-by-user
      [filename]
      (->> (lines filename)
           (parse-log)
           (code-357-by-user)))
    
    ; Throws IOException "Stream closed"
    #_(count-by-user "sample.log")
    
    
    
    ; Works, but I/O is coupled with the logic.
    
    (defn count-by-user
      [filename]
      (with-open [in (io/reader filename)]
        (->> (line-seq in)
             (parse-log)
             (doall)
             (code-357-by-user))))
    
    #_(count-by-user "sample.log")
    
    
    
    ; Separates out I/O. Allows us to compose the processing.
    
    (defn process-log
      [filename f]
      (with-open [in (io/reader filename)]
        (->> (line-seq in)
             (f))))
    
    ; Look at the first 10 lines that parsed
    #_(process-log "sample.log" #(->> % parse-log (take 10) doall))
    
    ; Count up all the "code 357" errors by user
    (defn count-by-user
      [filename]
      (process-log filename #(->> % parse-log code-357-by-user)))
    
    #_(count-by-user "sample.log")

    Log file sample:

    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | transaction failed while updating user joe: code 357
    2019-05-14 16:48:56 | process-Poster | INFO  | com.donutgram.poster | transaction failed while updating user sally: code 357
    2019-05-14 16:48:57 | process-Poster | INFO  | com.donutgram.poster | transaction failed while updating user joe: code 357

    Previous 1 7 8 9 10 11 12 Next

    Related Podcasts

    Reply All

    1

    Reply All Games & Hobbies
    Inside VR & AR

    2

    Inside VR & AR Gadgets
    Note to Self

    3

    Note to Self News
    BrainStuff

    4

    BrainStuff Natural Sciences
    This Week in Tech (Audio)

    5

    This Week in Tech (Audio) News
    Hands-On Tech (Audio)

    6

    Hands-On Tech (Audio) Technology
    footer-logo

    Contact Us

    Toll Free: 844-670-7747

    Links

    • Home
    • Top Charts
    • Networks
    • Apps
    • Independents Podcasts
    • Podcast Advertising
    • Podcast News
    • Contact Us
    • About Us
    • Analytics & Insights

    Stay Connected

      Privacy, Terms of Use & Our Code of Ethics Protecting Content Creators Copyrights