TopPodcast.com
Menu
  • Home
  • Top Charts
  • Top Networks
  • Top Apps
  • Top Independents
  • Top Podfluencers
  • Top Picks
    • Top Business Podcasts
    • Top True Crime Podcasts
    • Top Finance Podcasts
    • Top Comedy Podcasts
    • Top Music Podcasts
    • Top Womens Podcasts
    • Top Kids Podcasts
    • Top Sports Podcasts
    • Top News Podcasts
    • Top Tech Podcasts
    • Top Crypto Podcasts
    • Top Entrepreneurial Podcasts
    • Top Fantasy Sports Podcasts
    • Top Political Podcasts
    • Top Science Podcasts
    • Top Self Help Podcasts
    • Top Sports Betting Podcasts
    • Top Stocks Podcasts
  • Podcast News
  • About Us
  • Podcast Advertising
  • Contact
Not in our directory?
Add Show Here
Podcast Equipment
Center

toppodcastlogoOur TOPPODCAST Picks

  • Comedy
  • Crypto
  • Sports
  • News
  • Politics
  • True Crime
  • Business
  • Finance

Follow Us

toppodcastlogoStay Connected

    View Top 200 Chart
    Back to Rankings Page
    Technology

    Functional Design in Clojure

    Each week, we discuss a software design problem and how we might solve it using functional principles and the Clojure programming language.

    Advertise

    Copyright: © 2018-2019, Christoph Neumann and Nate Jones

    • Apple Podcasts
    • Google Play
    • Spotify

    Latest Episodes:
    Ep 029: Problem Unknown: Log Lines May 17, 2019
    Show notes

    Nate is dropped in the middle of a huge log file and hunts for the source of the errors.

    • We dig into the world of DonutGram.
    • "We are the devops."
    • Problem: we start seeing errors in the log.
    • "The kind of error that makes total sense to whoever wrote the log line."
    • (04:10) We want to use Clojure to characterize the errors.
    • Why not grep?
    • Well...we grep out the lines, count them, and we get 10,351 times. Is that a lot? Hard to say. We need more info.
    • (05:50) Developers think it's safe to ignore, but it's bugging us.
    • We want to come up with an incident count by user.
    • "Important thing we do a lot in the devops area: we see data and then we have to figure out, what does it mean?"
    • "What story is the data telling us?"
    • The "forensic data" versus "operating data" (Ep 022, 027)
    • "With forensic data, you're trying to save the state of the imprint of the known universe over time."
    • (07:40) The log file is the most common forensic data for applications.
    • To characterize, we need to manipulate the data at a higher level than grep.
    • First goal: make the data useful in Clojure
    • Get a nice sample from the production log, put it in a file, load up the REPL, and start writing code to parse it.
    • Need to parse out the data from the log file and turn it into something more structured.
    • (12:05) Let's make a function, parse-line that separates out all the common elements for each line:
      • timestamp
      • log level
      • thread name
      • package name
      • freeform "message"
    • "A message has arrived from the developers, lovingly bundled in a long log line."
    • (13:10) We can parse every line generically, but we need to parse the message and make sense of that.
    • "We want to lift up the unstructured data into structured data that we can map, filter, and reduce on."
    • We will have different kinds of log lines, so we need to detect them and parse appropriately.
    • We want to amplify the map to include the details for the specific kind of log line.
    • We'll use a kind field to identify which kind of log line it is.
    • There are two steps:
      1. recognize which kind of log line the "message" is for
      2. parse the data out of that message
    • "Maybe we'll have a podcast that's constantly about constantly. It's not just juxt."
    • (18:45) How do we do this concisely?
    • Let's use cond:
      • flatten out all the cases
      • code to detect kind (the condition)
      • code to parse the details (the expression)
    • Can use includes? in the condition to detect the kind and re-matches to parse.
    • (20:00) Why not use the regex to detect the kind too?
    • Can avoid writing the regex twice by using a def.
    • It's less code, but now we're running the regex twice: once in the condition and once to parse.
    • We can't capture the result of the test in cond. No "cond-let" macro.
    • We could write our own cond-let macro, but should we?
    • "When you feel like you should write a macro, you should step back and assess your current state of being. Am I tired? Did I have a bad day? Do I really need this?"
    • (24:05) New goals for our problem:
      • one regex literal
      • only run the regex once
      • keep the code that uses the matches close to the regex
    • Similar problem to "routes" for web endpoints: want the route definition next to the code that executes using the data for that route.
    • Feels like an index or table of options.
    • (25:20) Let's make a "table". A vector of regex and handler-code pairs.
    • We need a coffee mug: "Clojure. It's just a list."
    • The code can be a function literal that takes the matches and returns a map of parsed data.
    • Write a function parse-details which goes through each regex until one matches and then invokes the handler for that one. (See below.)
    • (30:15) Once we have higher-level data, it's straight forward to filter and group-by to get our user count.
    • Once again, the goal is to take unstructured data and turn it into structured data.
    • "You have just up-leveled unstructured information into a sequence of structured information."
    • Can slurp, split, map, filter, and then aggregate.
    • (32:10) What happens when we try to open a 10 GB log file?
    • Sounds like a problem for next week.
    • "Fixing production is always a problem for right now."
    • "If a server falls over and no one outside of the devops team knows about it, did the server really fall over?"
    • "Who watches the watchers?"

    Related episodes:

    • 020: Data Dessert
    • 022: Evidence of Attempted Posting
    • 027: Collected Context

    Clojure in this episode:

    • #""
    • re-matches
    • case
    • cond
    • if-let
    • ->>
    • slurp
    • filter
    • group-by
    • def
    • constantly
    • juxt
    • clojure.string/
      • includes?
      • split

    Related links:

    • GraalVM

    Code sample from this episode:

    (ns devops.week-01
      (:require
        [clojure.java.io :as io]
        [clojure.string :as string]
        ))
    
    ;; General parsing
    
    (def general-re #"(\d\d\d\d-\d\d-\d\d)\s+(\d\d:\d\d:\d\d)\s+\|\s+(\S+)\s+\|\s+(\S+)\s+\|\s+(\S+)\s+\|\s(.*)")
    
    (defn parse-line
      [line]
      (when-let [[whole dt tm thread-name level ns message] (re-matches general-re line)]
        {:raw/line whole
         :log/date dt
         :log/time tm
         :log/thread thread-name
         :log/level level
         :log/namespace ns
         :log/message message
         }))
    
    (defn general-parse
      [lines]
      (->> lines
           (map parse-line)
           (filter some?)))
    
    
    ;; Detailed parsing
    
    (def detail-specs
      [[#"transaction failed while updating user ([^:]+): code 357"
        (fn [[_whole user]] {:kind :code-357 :code-357/user user})]
       ])
    
    (defn try-detail-spec
      [message [re fn]]
      (when-some [matches (re-matches re message)]
        (fn matches)))
    
    (defn parse-details
      [entry]
      (let [{:keys [log/message]} entry]
        (if-some [extra (->> detail-specs
                             (map (partial try-detail-spec message))
                             (filter some?)
                             (first))]
          (merge entry extra)
          entry)))
    
    
    ;; Log analysis
    
    (defn lines
      [filename]
      (->> (slurp filename)
           (string/split-lines)))
    
    (defn summarize
      [filename calc]
      (->> (lines filename)
           (general-parse)
           (map parse-details)
           (calc)))
    
    
    ;; data summarizing
    
    (defn code-357-by-user
      [entries]
      (->> entries
           (filter #(= :code-357 (:kind %)))
           (map :code-357/user)
           (frequencies)))
    
    
    (comment
      (summarize "sample.log" code-357-by-user)
      )

    Log file sample:

    2019-05-14 16:48:55 | process-Poster | INFO  | com.donutgram.poster | transaction failed while updating user joe: code 357
    2019-05-14 16:48:56 | process-Poster | INFO  | com.donutgram.poster | transaction failed while updating user sally: code 357
    2019-05-14 16:48:57 | process-Poster | INFO  | com.donutgram.poster | transaction failed while updating user joe: code 357

    Ep 028: Fail Donut May 10, 2019
    Show notes

    Christoph has gigs of log data and he's looking to Clojure for some help.

    • Introducing a new topic.
    • The last few weeks we focused on Twitter and automatically posting to it.
    • Surprised by how much there is to talk about in a focused problem.
    • "There will always be more problems for the world to solve."
    • (01:53) Imagine if you will, the world of DonutGram!
    • A fictitious social network where people post their donut experiences.
    • This series is not about making DonutGram, but living DonutGram.
    • "Oh, the dark underbelly of application development: support."
    • "Let's take the shiny rock and flip it over and look at all the worms and bugs."
    • We talked about forensic data before, and much of DevOps is about looking at that data.
    • You want to paint the most complete story. It's a development problem too.
    • Most of the time, forensic data is written to a log file. It's the first line of investigation.
    • Imagine all of the components of your application speaking into one pipe, and you now get to reconstruct what happened.
    • (05:59) Everything is rosy at DonutGram, but then users start having issues.
    • Users get the "Fail Donut" and start posting that.
    • We will be assuming the role of the heroic DevOps team.
    • There is quite a bit of data written to the log file, more and more with each bugfix.
    • We're going to use Clojure to investigate this problem.
    • (08:50) What problems might we encounter?
    • Problem 1: Size
      • No one configured log rotation!
      • The log file is too large to load for analysis.
    • Problem 2: Unstructured data
      • Everything is a line or multiple lines.
      • We can build up layers of abstractions to gradually build understanding.
    • Problem 3: Non-linear data
      • We want to tell a story about what went wrong.
      • Many times, the pieces of that story are in different areas of the log file, and must be collated.
      • There can be multiple competing stories.
      • The parts of the story are like dominoes. How do you know if the third domino is important without knowing that the first two have fallen?
    • Problem 4: Alerts
      • "If you have a really good story to tell, how do you tell anyone about it?"
      • Alerts should be timely, depending on the audience.
    • Call out to the audience: do you have battle stories about processing log files? Let us know!

    Clojure in this episode:

    • nil

    Ep 027: Collected Context May 03, 2019
    Show notes

    Nate and Christoph reflect on what they learned during the Twitter series.

    • 6 months of podcast episodes!
    • Situated programs: Part of the real world. Affect and affected by the world around it.
    • (04:25) Concept 1: Was very helpful to focus on nailing down the algorithm before diving into the code.
    • (06:00) Concept 2: Two kinds of data: operating data and forensic data
      • Operating data: information you use for logic. Make it aggressively minimal.
      • Forensic data: information you record, off to the side, to tell a story later. Record as much as feasible.
      • Don't do logic on the forensic data!
      • If you end up writing code to sift through the forensic data in the future, cherry pick out the data need to use and treat that as operational data for that problem.
    • "'No' is temporary. 'Yes' is forever."
    • (09:45) Concept 3: Use components to wrap integrations
      • We wrapped DB and Twitter integrations.
      • If a component needs a resource, that resource should be a component.
      • Because resources are components, you can substitute the real component with a "faker" component.
    • (11:10) Concept 4: Fake resources for productivity
      • A "faker" component takes the place of a real component to enable developers to be more productive.
      • Faking is not for testing code, it's for effective development.
      • Can avoid Internet and latency.
      • Can run the code right out of the repo.
      • Allows you to explore how the whole system behaves during edge cases and behaviors.
      • Allows you to start with useful information or quickly reset back to scenarios.
    • "If an edge case is going to be annoying in the wild, you might as well get used to it being annoying during development."
    • (13:40) Concept 5: Aggressively decouple decision logic from imperative work
      • Push side effects (like I/O) to the "edges", but some situations make that less straightforward.
      • Challenge: our algorithm had a tick-tock: side effect, logic, side effect, logic, ...
      • Our approach: separate each side effect and only do one at a time.
    • (16:10) Concept 6: Watch out for convenient side effects
      • Eg. Doing a little bit of I/O in the middle of a function.
      • Eg. Get "now" from the system clock in the middle of a function
    • (16:35) Concept 7: Side-effect assumptions make your system fragile.
      • When you have code in one place that depends on something having been established by code in another place.
      • Eg. Knowing the calling function has already done a check.
      • There is an implicit umbrella of context. Creates a hidden dependency between those two places.
      • When the assumption is violated, there's nothing to stop the code making the assumption from just marching off the cliff.
      • Put the guard logic next to the code that needs it.
    • (19:15) Concept 8: Make context explicit
      • A function should take all of its "knowledge" in as data. It should not assume something has been done for it elsewhere.
      • "This forces you to write down all the things that matter instead of assuming something else has put things in place."
      • Explicit context allows you to separate out imperative steps.
      • "The hidden relationships, the ones that are only in my brain, those are the things that will completely burn me in 6 months."
    • (20:45) Concept 9: Favor instructions as data rather than instructions as function names.
      • Eg. Our Twitter wrapper only has one invoke function. (Inspired by aws-api.)
      • Function names aren't part of the data. If the "command" name is the function name, how do you match on it? You can't.
      • When commands = function names, you have implicit information: "I must have command X because I'm on line Y."
      • Observation: the description of the goal is usually more stable than the steps to achieve the goal.
      • If all of the instructions are in data, the processing component can completely change how it achieves the goal without the calling component changing how it interacts.
    • (26:00) Concept 10: Treat components within your application as loosely coupled services.
      • What is a "microservice"? When you make a request (aka "data"), ship it out, and get a response back (aka "data").
      • In Clojure, we can treat component interactions like loosely coupled APIs too: send them rich data, get rich data back.
      • Get loose coupling on the inside of the service, not just between services.
    • (27:25) We want to help programmer effectiveness and reduce burnout.
    • "We want to help you avoid situations where you're beating your head against the wall wondering where things went wrong."
    • "One of the great benefits of simple design and Clojure, it allows you create things that are fun to create AND fun to maintain."

    Message Queue discussion:

    • (29:15) "You know what else is a lot of fun? Hearing from our listeners!"
    • We had an issue with our SSL certs. Thanks for the heads up and let us know if you have any tech trouble.
    • "Automagically". We can't say we invented it, but we love that word.
    • "What is magic? An invisible force with a visible effect. Sound an awful lot like a side effect!"
    • "You start wielding the power, it feels GREAT! And then it all starts going wrong."
    • Easy does not imply simple. Go watch that talk. (See below.)

    Related episodes:

    • 020: Data Dessert
    • 021: Mutate the Internet
    • 022: Evidence of Attempted Posting
    • 023: Poster Child
    • 024: You Are Here, But Why?
    • 025: Fake Results, Real Speed
    • 026: One Call to Rule Them All

    Clojure in this episode:

    • nil

    Related links:

    • Cognitect aws-api
    • Simple Made Easy

    Ep 026: One Call to Rule Them All Apr 26, 2019
    Show notes

    Christoph thinks goals are data, not function names.

    • We were talking about the twitter handle again.
    • Last week, we talked about faking. It's not mocking.
    • The magic that makes it possible is using a protocol.
    • Switch out the real handle or the fake handle in component based on configuration.
    • "Yes, I do want to speak to the log file."
    • "Sometimes, the log file gets lonely."
    • (02:43) Christoph wants to talk about the protocol.
    • We made a function for each operation we did in Twitter.
      • Post tweet
      • Search
      • Get timeline
    • Each function took the information needed for that operation as individual parameters.
    • (04:32) Example: the AWS API wrapper that Cognitect released.
    • Other wrappers are enormous, with functions for each AWS functions.
    • Cognitect's wrapper has just one: invoke. At least only one that gets work done.
    • Unconventional (aka "weird").
    • Take a step back. What needs to happen to call a remote API?
    • We need to make an HTTP call to some end point.
    • "100% of the information that that server needs to get from us, we transmit as data."
    • "The path, the headers, everything about it is data."
    • "The function is in the URL and the URL is in the data. It's all data."
    • If you want the operations to be exposed as functions, you end up promoting the operation to be a function name, but when making the request you need to put the operation back into the data.
    • Benefit (not really): we get to write a lot of boilerplate code.
    • "Nothing like some boilerplate to get you warmed up in the morning."
    • (09:15) Back to the Twitter handle, what if we just had one function?
    • "Making a function for each endpoint goes against what we talked about last week, which is making our context explicit."
    • (09:55) First benefit of one function: Simplifies the worker half of our algorithm.
    • Before:
      • Decider creates map with {:command :twitter/fetch-timeline :twitter/last-tweet-id 1234}
      • Worker uses a multi-method to convert that to (handle/fetch-timeline handle {:twitter/last-tweet-id 1234})
      • Handle function will construct Twitter request {:url "https://twitter.com/fetch-timeline" :last-tweet 1234}
      • Handle sends request to Twitter and handles the response.
    • After:
      • Decider creates map with {:command :twitter/operation :twitter/command :fetch-timeline :twitter/last-tweet-id 1234}
      • Worker's multi-method detects a Twitter operation and passes the entire operation map to the handle.
      • Handle transforms the operation into the data needed for the request.
      • Handle sends request to Twitter and handles the response.
    • When the data is in the map, it's easier to test.
    • (18:03) Second benefit of one function: a place for common code in the handle.
    • Inside the handle, you can use a multi-method.
    • If there is common code, like auth checking or data transformation, that can go into the single function.
    • The same tick-tock that was used to get separate side effects from logic in our algorithm can be used inside the handle.
    • "Imperative logic is like a branching tree of side-effects."
    • The handle works on a rich map of information, and with one function, we're handing it one rich map of information.
    • "You have deciders all the way down."
    • (20:52) How do we do spec this single function?
    • One function needs to take data in multiple shapes. Making a spec that matches all of those shapes will be difficult.
    • This function call other, more specific functions to get its work done, so there can be specific specs on those.
    • "You don't have to spec everything at all the levels."
    • We can also lightly check the entry point to check for attributes that are common to all shapes.
    • "Spec is an open system. I can check for the presence of what I need without having to assert a closed world."
    • (22:55) Thinking about the symmetry. The following are equivalent:
      • 4 functions that each take 5 parameters.
      • 4 functions that each take a map with 5 keys.
      • 1 function that takes a map with 6 keys (the last one being the operation).
    • Programming systems that are strongly typed push you toward more functions and less data.
    • Trampoline from data to function to data to function to data...
    • Clojure's structural checking makes it so that we can have assurances about data without being forced to check or name everything.
    • "Say what you mean, don't make me read your body language and guess."

    Message Queue discussion:

    • (25:50) Separating logic from side effects.
    • "The tree of side effects is like a bowl where you swirl around into the bottom and then swirl back up to the top."
    • In a call stack, side-effects should be shallow on the stack.
    • "Pure things can't screw you over."
    • "Keep your friends close and your side-effects closer."
    • Running this concept out to its logical conclusion enabled us to discover the tick-tock approach.

    Related episodes:

    • 006: All Wrapped Up in Twitter
    • 021: Mutate the Internet
    • 022: Evidence of Attempted Posting
    • 023: Poster Child
    • 024: You Are Here, But Why?
    • 025: Fake Results, Real Speed

    Clojure in this episode:

    • defprotocol
    • defrecord

    Related links:

    • Cognitect aws-api
    • Stuart Halloway: Running with Scissors

    Ep 025: Fake Results, Real Speed Apr 19, 2019
    Show notes

    Nate wants to experiment with the UI, but Twitter keeps getting the results.

    • "This thing that we're making because we're lazy has now taken 4 or 5 weeks to implement."
    • Last week: "worker" logic vs "decider" logic. Allowed us to flatten the logic.
    • "You can spend months on the backend and people think you've barely got anything done. You can spend two days on the UI and then people think you're making progress."
    • (04:20) Now we want to work on the UI, but we don't want data to post to real Twitter.
    • "Why do you keep posting 'asdf asdf asdf' to Twitter?"
    • UI development involves a lot of exploration. Need to try things out with data.
    • We want to be able to control the data we're working with.
    • "We want carefully-curated abnormal data to test the edges of our UI."
    • (06:30) We could have a second Twitter account to fill with junk data.
    • Or, we could use the Twitter API sandbox.
    • Problem: we don't have much control over the data set. Eg. We can't just "reset" back to what it used to be.
    • Problem: what about when we're hacking on a plane?
    • Plus, we want to be able to share test data between developers.
    • (09:10) What can we do instead? Let's make a "fake" Twitter service we run on our machine.
    • "Fake" Twitter gives us something our application can talk to that is under our control.
    • Having a "fake" Twitter service creates new problems
      • Project wants to grow larger and larger
      • Now we have to run more things
      • Creates more obstacles between checking out the code and getting work done
    • (12:35) Rather than a "fake" Twitter service, we want to "fake it" inside the application.
    • What is "faking it"? Is this the same as "mocking"?
    • Use a protocol for our Twitter API "handle". Allows for an alternative implementation.
    • Not the same as mocking
      • We are not trying to recreate all the possibilities.
      • We are not trying to use it to test the component code.
    • The purpose is different than mocking. The Twitter "faker" is to help us work on the rest of the application.
    • The "faker" is about being productive, not about testing.
    • (17:20) Can have the faker start with useful data.
    • Default data in the faker can launch you straight into dev productivity.
    • "You want automation to support your human-centric exploration of things."
    • "If your environment is as fast as it can possibly be, then it allows you to be as fast as you can possibly be."
    • (20:10) Can use the faker for changing the interaction
    • Eg. have a "slow" mode that makes all the requests take longer.
    • Useful to answer the question: "What does this UI feel like when everything gets slow?"
    • "The faker can implement behavior you need, for exploring the space you need to cover, to converge on the right solution."
    • (22:00) Can have the faker pretend something was posted manually.
    • Allows you to see how the UI behaves when the backend discovers a manual post.
    • (23:00) Goals of faking
      • Support exploration. Not about testing and validation
      • Support the human, creative side of development
      • Support the developer experience

    Related episodes:

    • 006: All Wrapped Up in Twitter
    • 021: Mutate the Internet
    • 022: Evidence of Attempted Posting
    • 023: Poster Child
    • 024: You Are Here, But Why?

    Clojure in this episode:

    • defprotocol
    • defrecord
    • component/system-map

    Related projects:

    • Component

    Ep 024: You Are Here, but Why? Apr 12, 2019
    Show notes

    Christoph needs to test his logic, but he must pry it from the clutches of side effects.

    • Last week we ended up with "imperative mud": lots of nested I/O and logic.
    • (01:45) Difficult to test when all the I/O and logic are mixed together.
      • Object-oriented programming advocates for "mocking" objects with side-effects
      • Mocking reinforces and expands the complexity
    • Complexity stems from all the different branches of execution: mainline, exceptions, corner cases, etc.
    • With the imperative approach, no way to just "jump" to step 2 and test that in isolation. Have to setup all the mocks to tilt the flow of control in the desired direction.
    • The Go language handles the "error flow" problem by making error returns explicit (no exceptions), but that further compounds the branching problem.
    • (05:25) A different approach to deal with nesting: "happy path" route
      • Mainline of execution is the most common case.
      • Any error is thrown which has to be handled out of band.
    • "Exceptions are much better at mocking you than you are at mocking them."
    • Even more problems with mocking: suppose you want to test exponential backoff
      • API handle mock has to fail a couple of times and then succeed.
      • Requires a "stateful" mock!
    • "You have whole object hierarchies that are only in use for your tests!"
    • (07:35) Key observation: imperative programming uses the location of execution as a way of encoding information.
      • Code in the first branch of the "if" statement "knows" something different than code in the second branch.
      • Position in the flow of control is implicit information
    • Change implicit information to explicit information and you can flatten out the code.
    • Implicit: "You are here because you did X."
    • Explicit: "Separating where I am from what to do next."
    • "In an imperative flow, the fact you made it to line 3 means something."
    • (09:55) To get out of the problem of nesting, we need to have positive, represented information instead of implicit, contextual information.
    • Pattern is: logic, I/O, logic, I/O, logic, I/O, etc.
    • Capture each one as a multimethod: a "worker" method and a "decider" method.
      • The "worker" method dispatches on a "command", just does one thing, and returns the result.
      • "My job is not to question why. My job is but to do or die."
      • "I don't know if we need to get into the class struggle."
      • The "decider" method takes a command and a result, looks at them, decides what to do next, returns a command.
    • (18:50) You can work out all the scenarios, and the logic for each scenario is pure!
      • Testing becomes a matter of setting the data
      • Steps can be tested independently
    • (19:45) A "result" isn't simply spewing back the raw data from the I/O
      • Worker picks out the relevant parts and return those
      • Use spec or schema to document responses
      • Can use meta or a nested key (like :raw) to attach raw data for debugging or the "attempt" log (see Ep 022)
    • Logic should operate on well-known fields, but recording raw responses is useful for telling a story later on.
    • (22:45) Main flow of execution becomes a simple loop + recur
      1. Dispatch command to worker and bind response
      2. Dispatch command + response to decider and bind new-command
      3. (when new-command (recur new-command))
    • (24:30) A downside of the approach: it can be harder to see cases that are not covered, like an unsupported command or unhandled response
    • Once again, unit tests can help ensure coverage for the "deciders".
    • (27:30) How to test the worker?
      • They don't have any logic except doing the side effect.
      • Even "picking" the right values from the data could be factored as a pure function
      • Can test with the REPL. Try out the side effect and then leave it alone.
    • (29:50) Application testing creates another problem
    • What about testing the UI? We want to see changes in the data, but we don't want to actually post for real.
    • Great topic for next time.

    Message Queue discussion:

    • (32:05) Do we post the source code for episodes?
    • We post some code in the show notes, but less so for a high-level series.
    • Send us links to your code and we'll put them on the website.

    Related episodes:

    • 021: Mutate the Internet
    • 022: Evidence of Attempted Posting
    • 023: Poster Child

    Clojure in this episode:

    • loop
    • recur
    • meta
    • when

    Ep 023: Poster Child Apr 05, 2019
    Show notes

    Nate gets messy finding ingredients for his algorithm cake.

    • Last week we focused on how to determine what to post.
    • This week we focus on getting the data so we can compare it.
    • (01:55) Once again, we'll use component to organize our app.
    • What components should we have?
    • Component 1: The worker that wakes up, checks the DB, checks Twitter, and posts as necessary.
    • Debate: Do we need more than one component?
    • Question: What does that worker need? Those should be components.
    • (03:00) Component 2: The database connection.
      • Needs to be threadsafe
      • Allows all the DB logic in one place.
      • start method is a natural place to do migrations, indexing, etc.
    • (04:45) Aside: quick component refresher
      • Very lightweight mechanism for shared, stateful resources.
      • Dependency injection: components depend on other components, but framework instantiates them.
      • Each component has two lifecycle methods start and stop.
      • Setup state in start: initialize internal data, open connection, prep work, etc.
      • Tear it all down in stop: finishing work, closing connections, etc.
      • Declare all your components in a system.
    • (06:00) What should we call Component 1? "Poster"?
    • "It could be the poster child of components!"
    • (06:50) Component 3: The Twitter API handle
    • What is a "Twitter API handle" (aka the "Twitter handle")?
      • Data structure and code that handles authenticating and making requests.
      • Not a long standing socket connect (like the DB), but there is still state (auth key).
    • (08:50) What happens when the cached credentials expire?
    • Two options from Ep 006:
      1. Request function returns a tuple: [updated-handle, result]. The handle will only change when it has to re-auth.
      2. Use an atom for the handle. Request function mutates the handle when it has to re-auth.
    • In both cases, the Twitter wrapper deals with re-auth. The difference is whether the code that uses the handle needs to know about it.
    • We'll use #2 for our Twitter component.
    • We must still consider a race condition between components who all use an expired handle at the same time. Extra work will be done, but the result will still be correct.
    • If components use separate handles, they can't benefit from sharing the re-auth.
    • (12:55) We need a way to trigger "waking up"
    • Use at-at to schedule an interval for calling a function
    • Backed by Java's ScheduledThreadPoolExecutor.
    • "It's good to have your components be tidy and clean up after themselves."
    • Important to call at-at/stop in your component/stop method so component.repl/reset doesn't make more and more timers!
    • (16:10) We have the ingredients to implement the algorithm
    • Basic steps:
      • Wrap it all in an exception handler
      • Use a sequence in a let block to assign results back from Twitter and DB
      • Take results and pass them to pure function to determine what to do next
    • Complication: order dependent. We need last posted Tweet to know how far back to fetch from Twitter
    • New steps:
      • Fetch from DB: last posted and next scheduled
      • Use last posted ID to fetch from Twitter
      • Pass tweets and next scheduled tweet to pure function to determine if we should post
      • If we need to post, try posting to Twitter and capture result
      • If success, write tweet ID into database to mark as "completed"
      • No matter what, write the result to the "attempt" log
    • (22:00) Feels very imperative
    • "Do I/O. Do some logic. Do I/O. Do some logic. All those little bits of logic are very hard to test."
    • OO answer: mock the resources
    • Mocking makes more problems: now you have to implement all sorts of fake logic
    • "You just start grabbing more and more side effects and glomming them on to this big ball of mud, just so you can test a little bit of logic."
    • "Next thing you know, you're developing a whole vocabulary of mock creation."
    • We'll look at this more next week.

    Message Queue discussion:

    • (25:00) We were mentioned on the Illegal Argument podcast
    • Comparing Clojure REPL and Smalltalk REPL.
    • Common problem: state in the REPL does not reflect what is in the source.
    • For us, connected editor helps avoid that: we run pieces directly from our source files.
    • Still can end up out of sync: dangling symbol references.
    • Using tools.namespace.repl/refresh will find those dangling references.
    • Can still build up in pieces, but can use refresh to check it all at once.

    Related episodes:

    • Twitter handle, retrying on failure
      • 006: All Wrapped Up in Twitter

    Related projects:

    • Component
    • component.repl
    • At-At
    • Java ScheduledThreadPoolExecutor
    • tools.namespace

    Clojure in this episode:

    • atom
    • component/
      • start
      • stop
      • system-map
    • component.repl/
      • reset
    • at-at/
      • mk-pool
      • every
      • stop

    Ep 022: Evidence of Attempted Posting Mar 29, 2019
    Show notes

    Christoph questions his attempts to post to Twitter.

    • This week, continuing to dig into the "Twitter problem". We want to post to Twitter on a schedule.
    • "Writing code to help out with laziness."
    • Start with data to keep track of: inside (our data) and outside (Twitter data)
    • "Data from a foreign land."
    • We need to determine our "working view" of Twitter's data.
    • What is in our data? For each "scheduled tweet":
      • Text to post: the "status"
      • Timestamp of when to post
    • Timestamps are nice
      • Milliseconds since the epoch
      • Universal instant
      • Allows the client to localize
    • How do we know a scheduled tweet has been posted? A "posted?" boolean?
    • Boolean says, "Yes! It has been posted somewhere on the Internet."
    • Correlating identifiers are more useful than a Boolean.
    • The tweet ID is a correlating identifier. We can use it to lookup all of Twitter's data about it.
    • "We don't need to store all of Twitter in our database."
    • What is the story you need to tell about what happened?
      • A record of all the attempts allows us to tell a story about what happened.
      • Useful to have the timestamp of when our application posted it.
    • Make a separate log for attempts.
      • Attempting to post is a separate concern than what to post.
      • Don't complicate the scheduled tweet information by embedding the log.
    • "Once you have all the data, it allows you to ask new questions you didn't originally think of."
    • Clojure makes it easy to work with a large tree of data that came from an external source. We don't have to care about the structure of that data. We can just write it down.
    • Simply attempt to post the next scheduled tweet that does not have a Twitter ID recorded.
    • If it fails, just record the attempt, and go back to sleep.
    • "Handle the brick in front of you, and if you keep doing that, you'll eventually build the wall."
    • What if we don't hear the success response from Twitter, but it did get posted?
    • Idea: Try to detect if a tweet has already been posted.
    • If we can uniquely identify something by its content, we can know two things are the same without having a common ID.
    • Problem: Twitter can alter the contents.
    • Idea: fuzzy "measure of similarity" between our recent tweets and the next scheduled tweet.
    • We can record the fuzzy match in our attempt log too!
    • If we can correlate by contents, we could even identify when we manually post in advance.
    • As soon as you can determine equality by the substance of the thing itself, you can have more than one writer.
    • How "recent" is "recent"? Is it 100? Is it 200? Is it 500?
    • Even better, fetch all the tweets since the last ID we recorded.
      • we know we're seeing all of the tweets
      • can scan each of those for a match (in the case of a manual post)
      • know when the tweet stream ends, so we can know a posting is still needed
    • The worker will get there eventually. Can just give up on an error. No complex retry and recovery logic.
    • With more than one writer, we still can have a race condition. Ultimately Twitter has to deal with deduplication to avoid a double post in a short interval.

    Message Queue discussion:

    • Namespacing in a map is really useful
    • A flat, namespaced map is easier to traverse than a nested map.
    • One use for namespaces: indicate the origin of the data
    • Eg. :twitter/id, :twitter/status vs :local/id, :local/text
    • You see the namespace in your code, so it makes the data origin very visible.

    Related episodes:

    • Schemas for data. "Internal" vs "external" data.
      • 007: Input Overflow
    • Creating a "big bag of data" and asking it questions.
      • 018: Did I Work Late on Tuesday?
      • 019: Dazed by Weak Weeks
      • 020: Data Dessert

    Related projects:

    • XKCD: Is It Worth the Time?
    • The rsync algorithm

    Clojure in this episode:

    • pr-str

    Ep 021: Mutate the Internet Mar 22, 2019
    Show notes

    Nate wants to tweet regularly, so he asks Clojure for some help.

    • Problem: pre-author tweets so they can get posted automatically.
    • Want to make a "full stack" application this time.
    • Sounds complicated, what do we need?
      • Database of tweets: text to tweet and timestamp when it should be tweeted.
      • Frontend is a single-page application (SPA) that makes XHR ("AJAX") requests to the backend.
      • Needs to be able to wake up and post.
      • Persistent process backend.
    • We will not cover all these parts.
    • "You have a problem, so you make a UI to solve your problem. Now you have 2 problems." "More like 18 problems!"
    • We will focus on logic interacting with Twitter.
      • When should it post?
      • How does it know if a tweet has been posted?
      • What to do when Twitter returns an error?
    • Overarching theme: how do you deal with side-effects in a functional way?
    • Remember Ep 020: push side-effects and I/O to the edges.
    • Easy to fetch "current" time, but that's a side effect!
    • "Just because it's easy doesn't mean it's pure."
    • If you make time a parameter, all of a sudden you can mess with it!
    • "If only real time was a parameter we could manipulate."
    • Lots of fun to be had in the upcoming episodes.
    • "'Start with the data' is something we've come to again and again. If you can model the data, that's a very good place to start."

    Related episodes:

    • 006: All Wrapped Up in Twitter
    • 020: Data Dessert

    Related projects:

    • Maria

    Clojure in this episode:

    • nil

    Ep 020: Data Dessert Mar 15, 2019
    Show notes

    Christoph and Nate discuss the flavor of pure data.

    • "The reduction of the good stuff."
    • "We filter the points and reduce the good ones."
    • Concept 1: To use the power of Clojure core, you give it functions as the "vocabulary" to describe your data.
      • "predicate" function: produce truth values about your data
      • "view" or "extractor" function: returns a subset or calculated value from your data
      • "mapper" function: transforms your data into different data
      • "reduction" (or "reducer") function: combines your data together
    • Concept 2: Don't ignore the linguistic aspect of how you name your functions.
      • Reading the code can describe what it is doing.
      • Good naming is for humans. Clojure doesn't care.
    • Concept 3: Transform the data source into a big "bag" data that is true to structure and information of the source.
      • Source data describe the source information well and is not concerned with the processing aspects.
      • Transform into data that is useful for processing.
    • Concept 4: Using loop + recur for data transform is a code smell.
      • Not composable: encourages shoving everything together in one place.
      • "End up with a ball of mud instead of a bag of data you can sift through."
      • "You know what mud sticks to really well? More mud! It's very cohesive! And what couldn't be better than cohesive programs!"
    • Concept 5: Use loop + recur for recursion or blocking operations (like core.async)
      • Data shows up asynchronously
      • Useful when logic is more naturally expressed as recursion than filter + map + reduce.
    • Concept 6: Duality: stepwise vs aggregate
      • Stepwise problem: advance a game state, apply async event, stream processing, etc.
      • Stepwise: reduce, loop + recur
      • Aggregate problem: selecting the right data and combining it together.
      • Aggregate: filter + map + reduce
      • Aggregate problems tend to be eager--they want to process the whole data set.
    • Concept 7: Use your bag of granular data to work toward a bag of higher-level data.
      • We went from lines → entries → days → weeks
      • "Each level of data allows you to answer different questions."
    • Concept 8: Duality: higher-level data vs granular data with lots of dimensions
      • Eg. having a single "day" record vs a bunch of "entry" records that all share the same "date" field.
      • The "right" choice depends on your usage pattern.
      • Dimensional data tends to stay flat, but high-level data tends toward nesting.
      • A high-level record is a pre-calculated answer you can use over and over quickly.
      • Highly-dimensional, granular record allows you to "ask" questions spanning arbitrary dimensions. Eg. "What weeknights in January did I work past midnight?"
    • Concept 9: Keep it pure. Avoid side effects as much as possible.
      • Pure functions are the bedrock of functional programming.
      • REPL and unit test friendly.
      • "You can use data without hidden attachments. You remember side effects when you're writing them, but you don't remember them three months later."
    • Concept 10: Keep I/O at the "edges" with pure functions in the "middle".
      • "I/O should be performed by functions that you didn't write."
      • Use pure functions to format your data so you only have to hand it off to the I/O function. Eg. Create a list of "line" strings to emit with (run! println lines).
      • You can describe your I/O operations in data and make a "boring" function that just follows them. This allows you to unit test the complicated logic that determines the operations.
      • Separates out I/O specific problems from business logic problem: eg. retries, I/O exceptions, etc.

    Related episodes:

    • 002: Tic-Tac-Toe, State in a Row
    • 012: Fiddle with the REPL
    • 018: Did I Work Late on Tuesday?
    • 019: Dazed by Weak Weeks

    Clojure in this episode:

    • filter, map, reduce
    • loop, recur
    • group-by
    • run!
    • println

    Previous 1 8 9 10 11 12 Next

    Related Podcasts

    Reply All

    1

    Reply All Games & Hobbies
    Inside VR & AR

    2

    Inside VR & AR Gadgets
    Note to Self

    3

    Note to Self News
    BrainStuff

    4

    BrainStuff Natural Sciences
    This Week in Tech (Audio)

    5

    This Week in Tech (Audio) News
    Hands-On Tech (Audio)

    6

    Hands-On Tech (Audio) Technology
    footer-logo

    Contact Us

    Toll Free: 844-670-7747

    Links

    • Home
    • Top Charts
    • Networks
    • Apps
    • Independents Podcasts
    • Podcast Advertising
    • Podcast News
    • Contact Us
    • About Us
    • Analytics & Insights

    Stay Connected

      Privacy, Terms of Use & Our Code of Ethics Protecting Content Creators Copyrights