TopPodcast.com
Menu
  • Home
  • Top Charts
  • Top Networks
  • Top Apps
  • Top Independents
  • Top Podfluencers
  • Top Picks
    • Top Business Podcasts
    • Top True Crime Podcasts
    • Top Finance Podcasts
    • Top Comedy Podcasts
    • Top Music Podcasts
    • Top Womens Podcasts
    • Top Kids Podcasts
    • Top Sports Podcasts
    • Top News Podcasts
    • Top Tech Podcasts
    • Top Crypto Podcasts
    • Top Entrepreneurial Podcasts
    • Top Fantasy Sports Podcasts
    • Top Political Podcasts
    • Top Science Podcasts
    • Top Self Help Podcasts
    • Top Sports Betting Podcasts
    • Top Stocks Podcasts
  • Podcast News
  • About Us
  • Podcast Advertising
  • Contact
Not in our directory?
Add Show Here
Podcast Equipment
Center

toppodcastlogoOur TOPPODCAST Picks

  • Comedy
  • Crypto
  • Sports
  • News
  • Politics
  • True Crime
  • Business
  • Finance

Follow Us

toppodcastlogoStay Connected

    View Top 200 Chart
    Back to Rankings Page
    Technology

    On-Call Me Maybe

    A podcast about DevOps, SRE, Observability, On-Call, and everything in between.

    Advertise

    Copyright: © 2019 On-Call Me Maybe

    • Apple Podcasts
    • Google Play
    • Spotify

    Latest Episodes:
    From SCUBA to Kubernetes with Abby Bangser of Syntasso Nov 15, 2022
    Show notes

    About the guest:Abby Bangser is a UK-based Principal Engineer at Syntasso, delivering Kratix, an open-source cloud-native framework for building internal platforms on Kubernetes. Her keen interest in supporting internal development comes from over a decade of experience in consulting and product delivery roles across platforms, site reliability, and quality engineering.Abby is an international keynote speaker, co-host of the #CoffeeOps London meetup, and supports SLOConf as a global captain. Outside of work, Abby spoils her pup Zino and enjoys playing team sports.Find our guest on:Abby’s TwitterAbby’s LinkedInAbby’s MastodonFind us on:On Call Me Maybe Podcast TwitterAdriana’s TwitterAdriana’s MastodonAdriana’s LinkedInAdriana’s InstagramAna’s TwitterAna’s LinkedInAna's InstagramShow Links:SyntassoKratixKubernetesO11ycastClean CodeThoughtWorksQuality Assurance (QA)Parveen Khan on OCMMTracetestKubernetes ControllerAdditional Links:O11ycast Podcast: Ep. 16, Observability and Test Engineers with Abby Bangser of MOO#CoffeeOps London MeetupGlobal SLOConf CaptainAbby at Agile Testing Days 2022Video: Observability in Testing with Abby BangserMinistry of Testing: Meet the Instructor Podcast with Abby BangserSlight Reliability Episode 24 - Interview with Abby BangserTranscript:ADRIANA: Hey, y'all. Welcome to On-Call Me Maybe, the podcast about DevOps, SRE, observability principles, on-call, and everything in between. I am your host, Adriana Villela, with my awesome co-host...ANA: Ana Margarita Medina.ADRIANA: And today, we are talking to Abby Bangser of Syntasso. So welcome to the show, Abby. ABBY: Thank you so much for having me. ADRIANA: Awesome. Well, we're super stoked to have you. Ana and I were saying that we've been Twitter-fangirling you. [laughter]ABBY: All around with this crew. All around with this crewADRIANA: So yeah, we're super stoked to have you. First things first, what are you drinking today?ABBY: Yes. So I've brought along one of my favorite beers, the zero, the non-alcoholic version of a blanc, which is a specialty beer that is from 1664, but they only sell it in France. So all we get is the 1664 lager here in the UK. And so we took a trip to France and brought back some cases.ADRIANA: That's awesome. Then you're going to have to return to France to get more.ABBY: Oh, it's an annual pilgrimage to pick these up for sure. [laughs]ADRIANA: That is awesome. My drink today is a homemade bubble tea sans bubbles. I put basil seeds in mine, and I added some mango juice, so I got a little bit of flair.ANA: Both of those sounds so refreshing, and I'm so jealous because I'm just sipping on a classic Coca-Cola ice drink, and I'm just like, hmm, something's missing. I need a little bit of different flavor or just like even some boba. Can I do bubble tea with Coca-Cola? ADRIANA: Oh my God. [laughter] That's so cool.ABBY: I feel like that's the chaotic something in those quadrants you put it together. ANA: Chaotic evil.ABBY: Yeah, chaotic evil. [laughs]ADRIANA: That would be some really cool conference swag. I mean, it's perishable, but wouldn't that be neat?[laughter]ANA: I feel like we can do some interesting stuff with sodas and bubble tea at conferences that folks are not doing. I think, in general, I would love to see more tech conferences have non-alcoholic options. Just going to events where it's constantly just pushing booze, it's like, wait, that's not the most inclusive space. Like, you don't know where folks are at or the type of environment that it could do. So it's always really nice when I got to go somewhere, and we have those 0% lager showing up now more. or you're not pushing alcohol is like the main consumption that it's like, we have cold beverages. Just do things like that or offer mocktails too.ADRIANA: Yeah, totally, totally. I hear that the mocktail movement is growing. ABBY: The non-alcoholic beer is so good these days as well, beyond just the basic lagers. I find that when you just end up with the one basic lager that's not alcoholic, it's kind of a cop-out. And I've definitely pushed for more non-alcoholic options in places before. ADRIANA: Awesome. On to more techy things. Abby, why don't you tell us a little bit about the work that you do?ABBY: Yeah, absolutely. So as you mentioned, I work at a company called Syntasso. And what we are building is an open-source tool called Kratix, which is trying to help platform engineering teams build platforms. So there are lots of tools for how to deliver platform aspects, so deliver a database to software engineering teams that need one or deploy applications in an effective way. But how do you actually create a coherent story of a product of what your platform team delivers is the kind of problem space that we're trying to tackle. So it's been really interesting. It's my first-time, full-time engineering on Kubernetes. So we're building a controller to do that. And so it's been really fun getting to know that aspect of development.ADRIANA: That's awesome. You've had kind of a varied career. ABBY: [laughs]ADRIANA: I mean, I first heard about you from Ana when I was doing some research for a blog post on observability and testing. So that's where I first heard of you and the work that you were doing, and I heard your interview on o11ycast as well, which I absolutely loved. Yeah, so why don't you tell us a little bit about your career path, how you got to where you are now?ABBY: Oh, that I do need a beer for. [laughter] So I graduated at a time when it wasn't super obvious or easy to find jobs. So I graduated university in 2008. And so there were opportunities, but they were a little bit less than maybe when I started uni. And I didn't really know what I wanted to be when I grew up anyways. I was actually working as a scuba dive master at the time.ADRIANA: Oh.ABBY: And I figured I probably had to move on a little bit, wasn't sure to where. And I actually started working for an investment firm just doing data entry. And it was a really small firm. It was just three full-time people and then some data entry people who came in temporary. And they brought me on full-time to manage that data collection and analysis side of things. In doing so, it was actually generating the data we use to make smart investments, and that meant a lot of times scraping websites about opportunities that we had. So it was about real estate investment and finding information about the real estate that we were looking at. And so I started using a scripting tool that, I mean, this is remember 2008. This was basically Selenium written in VBScript but a proprietary software that I installed from a CD. I'm still quite young, but I still have some stories that make you wonder about tech. [laughter]ADRIANA: Holy cow.ABBY: It made me understand clean code before I ever knew the terms. I started all of a sudden realizing that I could reuse a script from one website to another if I did a good job of kind of isolating what was unique about that website and all these kinds of aspects of it, and I really enjoyed that. And a friend of mine worked at Thoughtworks in legal. And when I was saying I'm enjoying the coding side more than the investment side, she said, "Hey, well, we teach people how to be developers." And I showed up not knowing anything about Thoughtworks and how great it is with some code I'd written in a text editor and after reading Head First Java for a few days. And they kind of looked at me and said, "That's interesting. Probably not quite right for a developer job just yet, but really like how you think. Come join us as a QA." And what makes that really ironic is I then spent the next seven years working with them, absolutely pushing against the idea that people who failed out of developer interviews should go into the QA track despite that being how I entered the QA track. Because I felt like if they had put me through the interview for QA, I would have smashed it, and then I would have earned my position in QA rather than it being just, oh, you didn't quite cut it at dev, well, there's this other role. And I like to give them the benefit of the doubt that they knew I'd smash it, and that's why they pushed me that way. But I felt like it was important that you do that process. So yeah, I worked at Thoughtworks as a QA for seven years, moving from quite automation-focused to more analysis-focused to more DevOps and delivery-focused because if it works on your machine, does that really matter? Got to get to production. And then eventually into production systems and infrastructure-focused and then moved on from there to be in-house. I was really excited to get into a product, and I worked on a platform engineering team as a QA there and lead engineer for a couple of years and then as an SRE for a couple of years. And now, I find myself with my first title as a software developer. ADRIANA: That is so cool. ABBY: What a journey. Sorry, it's a bit of a mouthful there. [laughs]ANA: No, I love the journey because I think that's the beauty of technology. The space is so large, and there are so many different components that come into technology that until you touch something and you learn about it, you have that like, oh, this was really cool moment. But then, until someone is able to guide you in a way to see how everything works, even your transition from QA to going closer to reliability and DevOps was like, oh, I'm getting closer to things are working properly. But what is that business level, big picture of it as a business, like, what they really need? And then getting a chance to move on to SRE is just always really cool. I do have a question for you because something that we constantly talk about is just the terms always changing, and I'm really curious what your take is. Now that you're an engineer working on a platform that is building platform as a service, and you've been an SRE, and you've been an engineer on a platform team, how are things different? For folks that are always still trying to answer that question, is SRE supposed to be responsible for a platform, or is this something that's been held differently?ABBY: Oh, that's a tough one. And I think I have to fall back to the it depends, I'm so sorry. [laughter] But it does in that I think that it really...I'll go with a lesson I learned. So I actually was chasing the title SRE for a while. As I was getting more into the quality of the system rather than the quality of a feature or an application, I realized that doing that with the title QA could put up barriers that I didn't think were fair, but they existed. And so I thought if I can switch titles a bit, I can maybe help shuffle some of that responsibility and the opportunity back towards the title QA, but I need to break down those barriers first. So I'm not embarrassed to say I was chasing a title for a while. But when I got to that title, I think that it was very interesting to see how that plays out. SRE is different at different scale or organizations, at different types of organizations, different cultures of organizations. And in my experience, I was working as the title SRE in an organization that was quite small, and its problems were significantly more about how do we have a reliable, up-to-date database that has the correct backups and disaster recovery and things like that than it was around setting service-level objectives for our very small user base. And if I were to place where I am on the...I think of SRE as a bit of a spectrum of very deeply technically skilled in an area that can help with those nuance issues of resilience and reliability for a specific tech through to kind of the more higher level customer-focused service-level objective side of things, and I'm probably closer to that side. So this was maybe a misfit for what I was hoping to work on. Even if I can look at and go, yes, I see that as viable and reasonable SRE work, it's not the side of SRE work that I'm most interested by. My biggest learning is don't worry so much about titles; worry more about what are you getting to work on.ADRIANA: Yeah, that's so important. I think I spent so much of my career lamenting the fact that people my age were having these fancy titles and stuff, and I'm like, oh my God, I'm failing at life. And then I'm like, wait, but if I get to do awesome work that fulfills me, then that's what matters, right? ABBY: And that's actually how I got to learn as much as I have about platform engineering and things is the job I joined after Thoughtworks. I was looking at a role there. The manager there I knew and I knew from the QA community. And she said, "Look, we have this role open for QA. We aren't really actively interviewing for it because it's quite a niche role. And we don't really just want generic I write Selenium tests testers. We want people who think more globally about quality and things like that." And she was like, "You'd be perfect for it." And I was like, "Oh, but QA. [laughter] I'm really nervous about the title because I've seen...I've had brick walls after brick wall, and I busted through a lot of them. But I'm getting pretty tired. I'd like to be able to get into observability, and telemetry, and all these things." And she's like, "I can bring you in as a platform engineer, but your salary is going to reflect your experience there. I think that your experience as a QA puts you in senior bracket, and you get your senior salary, and you'll have that senior-level impact on the team. But I guarantee you that this team will not hold you back based on the title of QA." And I sort of just had to take a leap of faith with someone that I trusted, and it worked out brilliantly. So I think sometimes shying away from your titles can actually cause problems as well because they do open doors to more senior roles where you can have more influence over what it looks like. ADRIANA: Yeah, absolutely. And also having somebody who you know has your back and your interests at heart. I think that makes such a huge difference because then it makes you feel like you can contribute to the job and really put everything into it, right? ABBY: Absolutely.ANA: I mean, especially being transparent on the salary aspect of it where it's like, this is something that's going to really matter to you because it's how you live your day-to-day. Let's make sure that you're getting what you're worth, like, your years do matter. What do you say to folks that continue having this misconception about QA work? I mean, you did extensive work in QA. But I know that the industry still needs a lot of work for QA to get uplifted. Because we had one of our other guests that got to talk a little bit more about how they use observability in QA, and it's like, it's really extensive engineering work of understanding a system and explaining it to someone else. That work is going to possibly get the product engineer promoted, but the underlying work is being done by someone else that's going to not get credit.ABBY: There's such amazing work being done in QA. Nothing about what I said should be interpreted as me running away from a terrible part of the industry in any way. It's more about how other people perceive things and how other people open doors for you. The challenge is that every role in this industry can feel quite siloed. If you're in the SRE role and you're going to SREcon, and you're going to like SLOconf, and you're going to these…

    Full show notes at the publisher

    Kube Cuddles with Rich Burroughs of Loft Labs Nov 08, 2022
    Show notes

    About the guest:Rich Burroughs is a Staff Developer Advocate at Loft Labs where he's focused on improving the happiness of teams using Kubernetes. He's the creator and host of the Kube Cuddle podcast where he interviews members of the Kubernetes community. Rich was one of the founding organizers of DevOpsDays Portland, and he's helped organize other community events. He also has a strong interest in how working in tech impacts mental health. Rich has ADHD and has documented his journey on Twitter since being diagnosed.Find our guest on:Rich’s TwitterRich’s LinkedInFind us on:On Call Me Maybe Podcast TwitterAdriana’s TwitterAdriana’s LinkedInAdriana’s InstagramAna’s TwitterAna’s LinkedInAna's InstagramShow Links:Loft LabsLukas Gentele (Loft Labs CEO and co-founder)Fabian Kramm (Loft Labs CTO and co-founder)Kube Cuddle podcastKubernetesVclusterJoe Beda (Kubernetes Co-Creator)Craig McLuckie (Kubernetes Co-Creator)Ian ColdwaterKelsey Hightower TetrisHolly Cummins talks about zombie serversCustom Resource Definition (CRD)PostgresSQLiteElastic Kubernetes Service (EKS)CodefreshGitOpsKubernetes ReplicaSetREST APIMatt LemayAgile ManifestoSAFeADHDTranscript:ADRIANA: Hey, y'all. Welcome to On-Call Me Maybe, the podcast about DevOps, SRE, observability principles, on-call, and everything in between. I am your host, Adriana Villela, with my awesome co-host...ANA: Ana Margarita Medina. ADRIANA: And today, we are talking to Rich Burroughs, who is a Staff Developer Advocate at Loft Labs. Rich, welcome.RICH: Thanks so much. I'm really excited to be here. People may not know this, but Ana and I actually worked together at a previous job on the same team. So it's a pleasure to be talking with you both.ANA: We just happen to be bringing SRE friends together to talk about this amazing space and ways that we can make it better, the cool work that we're doing. So it's just an honor to bring you on board and share some of your learnings in the past few years. RICH: Well, wow. I thought everything was going perfect. I didn't know we needed to improve things.[laughter]ANA: You mean you don't constantly reconsider every choice you make in life, from what you're wearing to why you have such photos on your Twitter or why you retweet ADHD memes? [laughs]RICH: Yeah, I do still carry a decent amount of impostor syndrome, so...ANA: [laughs] It's fair. I think we all do, and it gets better, and it doesn't get better. That's my words of advice.RICH: [laughs]ADRIANA: I think if you're good at your job, you do have some form of imposter syndrome. So I feel like it's a gut check that we're doing something right.RICH: You know, it's funny; I was actually just tweeting about this a little bit ago. Somebody had mentioned the idea of people assuming you should know something that you don't know. And I work in the Kubernetes space, and I interact with a ton of people. And I bet that a lot of people would be shocked if they were to give me a Kubernetes quiz at how bad I would do. [laughter]ADRIANA: I feel your pain. [laughs] ANA: To be fair, the ecosystem of Kubernetes is just constantly changing, and it's amazing to see such large projects move. But thinking four years ago, where I was leading Kubernetes 101 classes and letting folks get up and running from scratch on Ubuntu servers, looking at an hour, a lot of it is just manage Kubernetes. And then we were just trying to focus more on some resource consumption, and networking, and security aspects of stuff. I'm like, yeah, [laughs] I can't be doing these workshops anymore. I don't feel like I have this extensive knowledge down in the stack.RICH: I mean, it's so complex. And that's a little bit of almost a trope at this point, you know, how complex Kubernetes is. But the reality is that there are all those little niches like security, and storage, and networking. You end up following people on Twitter or seeing their content or whatever who are experts in one or more of those areas. And I know in my brain sometimes I compare my level of knowledge with those people. And it's just not fair to me to compare my level of Kubernetes security knowledge with Ian Coldwater. Or I follow Joe and Craig; I mean, they invented Kubernetes, so obviously [laughs] they're going to know more about Kubernetes than I do. So I think that's part of it, too, is that, to me, learning and growth is about your own personal progression. And the important part is to try not to compare yourself too much with other people. But that's a lot easier said than done sometimes. ADRIANA: That's so true. One thing that I was wondering: how did you get into Kubernetes?RICH: Well, I guess, just let me back up a little bit from where I was going to start. I have a long background or had a long background in operations. So I started as a sysadmin in like 1995 or something. And so I had worked with Linux for many years. I did lots of different kinds of ops roles. And in 2015, I was at this small conference here in Portland that, like most people, probably never heard that it was even happening. It was very much a local event. And this guy named Kelsey Hightower was there. And he was working at CoreOS at the time, and he gave this talk. And you can still find versions of this talk online if you Google Kelsey Hightower Tetris, where he was playing Tetris during the talk and using that as a metaphor for Kubernetes. And the idea being that these compute nodes that we no longer care about which nodes our apps are running on or things like that. That suddenly, these compute nodes are just a bunch of memory, and CPU, and storage. They're just real sources that are getting consumed. And that was the way to think of them instead of thinking of them as the host that the front-end app runs on, you know, which was very much my background. And I sort of fell into this niche fairly early on in my career where I was...I wouldn't call it a systems administrator. That was my title, but I don't think this position even exists anymore, but we used to call it more of like an application administrator. So I was doing manual deployments of applications, and managing their configurations, and troubleshooting problems at the app level of the stack. And so that's a lot of where my focus was. And so, so much of what Kelsey was talking about really spoke to me because I was that person in my shop who knew which services run on which hosts. I could tell you right away, oh, the front-end service runs on this host. And I dealt with so many of those things he was talking about and had felt a lot of pain. And so the idea that there was this platform that took a lot of these practices that we already were doing as ops people and just kind of built them into the platform, the scheduling and all of that, I just thought it was brilliant.And I was hooked pretty much immediately but not really like a Kubernetes practitioner. Some people might be surprised to hear this, but I've actually not worked in a shop where we run Kubernetes as a main platform. So a lot of my experience with that over the years was just following along with the project and playing with it. But it was always something I was very, very interested in. And then, actually, I guess this would have been early 2020 or, no, late 2019. I had gone to a couple of KubeCons and really enjoyed that. And I went to the one in San Diego, and I was looking around and realized that I knew all these rad people in the community that I could have access to. And I had done some podcasting before and really enjoyed it. And I suddenly was like, oh, I could do a Kubernetes podcast. And a lot of podcasting or any kind of media stuff is getting access to people. And this is something that I think people don't think about necessarily. But if you're going to have a podcast, you have to have guests, right? [laughs] And so you've got to know some people or be able to get people to come on. And it just struck me that I knew a bunch of pretty influential people in the community and could probably get them to come on the show. So I started doing that. And then, about two months into the podcast, the lockdown happened. [laughs] And I was like, oh wow, this is the worst possible time in the world to launch a new project because now my mental health is in the garbage. And I'm just lucky to even be able to take a shower and get dressed, let alone try to sustain a project. But I managed to keep it going over time. And then now I'm working for my first time at an actual Kubernetes vendor, Loft Labs. So I've been there since April of 2021. So I've been really enjoying that. I just love this community. There are a lot of fantastic people in the Kubernetes community, a lot of super smart but also very, very generous and people with good values. Yeah, it's really great to be working with it for a living now.ANA: It's kind of awesome because it's that, it's like when we think about the work of Kubernetes specifically, we are seeing just in general, the way that we've been doing infrastructure is actually getting revamped. And there are these things that are like acknowledging that the systems are so complex and that we can't do things the same way. So we get to start this transformation. And going to that little piece where we're talking about impostor syndrome, a portion of it is also understanding that because it is a project that is so large, it's going to be 100% fair for you to only know one vertical or to just kind of be like, I contribute to this open-source project but in very different ways. Like, I just spent the last three months being part of the version 1.25 release team as a communication shadow. And it was very interesting to see all the portions that happen in order to get version 1.26 out the door. And I was like, never did I know that all these little things need to be actually checked in on every two or three days. And it takes as many people to get it done for people to then have 1.26 and then providers to start adopting it. So sometimes you think that you're not making any contributions, but you really are, and sharing those stories is huge. It's like, it's what makes us learn a new topic, or find a new mentor, or even feel like we belong in a community.RICH: This is another topic that came up not too long ago on Twitter. So I actually had my first pull request accepted into the Kubernetes project. And it's sort of a funny pull request because, in the Kubernetes project, there is a YAML file that controls the list of the channels that are in the Kubernetes community Slack. And we wanted to get a channel added for one of these open-source tools that I work with. And so I put in the pull request to do that, and the PR got accepted and merged finally. And so I was like, wow, I'm technically a Kubernetes contributor because I've got this PR merged. And I was talking about that on Twitter. And some people were like, "Well, no, you already were a contributor. [laughter] You're doing this podcast, and you're helping people who are using Kubernetes. And so you've already contributed a lot." And that's actually the point of view I have in general. And it's kind of funny because if I were talking to someone else, I would have said the exact same thing. But when it came to myself, I wouldn't give myself that credit which is interesting.ANA: I can relate. [laughs] It's like we are so lenient and have empathy for others. But sometimes, we forget that empathy starts with ourselves, where we allow for failure to happen. And we also sit down and introspect and ask ourselves questions of what we want to do or why we're doing something, just like get all aligned.RICH: Yeah, I can be very, very hard on myself.ANA: I wanted to ask for folks listening; what exactly does Loft Labs do?RICH: We're focused on Kubernetes multi-tenancy and self-service. So our commercial product, Loft, gives platform engineers a tool that they can use to give developers self-service access to Kubernetes environments. For me, I've worked in the past in roles where, like I said, I was working very closely with engineers deploying these apps and things. And there were times where the engineer had opened a ticket up for me, and I couldn't get to it for three days, and they were totally blocked. And I've been on the other side of that too, where I needed someone from...I needed a load balancer form setup, or I needed a port open on the firewall or something like that. And I was sitting around waiting for another team to do that for me. And so I very much felt the pain of being both the person waiting for someone to do something for them and also the person who feels kind of guilty because you've got this open ticket and you know somebody's blocked, but you just have other higher priority things that you need to be working on, you know. So I'm very much a believer in self-service. And when I first saw the product, that was one of the things that I thought was so important. And then I think that, like in terms of the multi-tenancy stuff, the thing that the platform has built into it is this concept of virtual Kubernetes clusters, which is a new way to share a Kubernetes cluster. To talk about it on a very high level, the idea is that you've got this cluster, and people tend to go one of two ways when it comes to provisioning clusters for tenants, for teams, either they do namespace isolation where they take one cluster, and they carve it up between a bunch of tenants and give them all a namespace. And that can work well in some scenarios, but it's got some problems too. Say that this is a dev cluster, and I'm a developer, and I want to be able to make CRDs that go along with my app. Well, as a normal tenant in a namespace isolated cluster, I'm not going to have access to those global objects like CRDs. So that's a problem, along with the fact that things just get really complex when you have to put in all these exceptions and network policies. Say somebody needs to have three namespaces, and they need them to talk to each other. There are all these things that can come up that make it more complex. And because of that, because it's hard to do, a lot of people default to the other option, which is the Oprah thing, you know, look under your chair, and everybody gets a cluster, [laughter], and that's just a nightmare. Like, from a management perspective, if you've got thousands of Kubernetes clusters lying around, besides the fact that it's expensive, it's like, how do you know that they're secure? How do you know what's running on them? How do you know that they're even needed anymore? Holly Cummins did a really great talk at KubeCon a few years ago about this. And she used the phrase zombie clusters for these clusters that are out there, and they've got workloads running on them, but they're not needed anymore. Like, nobody's actively using them. And the reality is that that has an actual impact on our environment because of the power and resources that are being used to keep these workloads running that don't even need to be running in the first place.So it's like that meme of the guy who's sweating, and he's got the two buttons to push. And it's like [laughter] one is the namespace isolation, and the other one is giving everybody a cluster. And I had heard about this for years from people in the community. I'd heard about the pain that people feel with multi-tenancy. And so when Lukas, our CEO, approached me about potentially working with him, I took a look at the tool, and I was like, you know…

    Full show notes at the publisher

    Observability Internships with Mohammad Harun, Software Engineering Student at McMaster University Nov 01, 2022
    Show notes

    About the guest:Mohammad Harun is a 4th-year software engineering student at McMaster University in Hamilton, Ontario, Canada. He spent this past year working as an Observability intern at Wavelo, helping to develop best practices around Observability at the company.Find our guest on:Mohammad’s LinkedInFind us on:On Call Me Maybe Podcast TwitterOn Call Me Maybe Podcast LinkedIn PageAdriana’s TwitterAdriana’s LinkedInAdriana’s InstagramAna’s TwitterAna’s LinkedInAna's InstagramShow Links:WaveloObservabilityOpenTelemetryMcMaster UniversityTranscript:ADRIANA: Hey, everyone. Welcome to On-Call Me Maybe, the podcast about DevOps, SRE, observability, on-call, and anything in between. I am your host, Adriana Villela, with my awesome co-host...ANA: Ana Margarita Medina.ADRIANA: And we are here today with Mohammad Harun of Tucows/Wavelo. He is a junior observability engineer. And actually, we've got a bit of a connection because I hired Mohammad at Tucows back when I was working there as my observability intern. So, welcome.MOHAMMAD: It's great to be here. Thank you, guys, for having me.ADRIANA: We're super excited to talk to you. So I guess the first question is, what are you drinking today?MOHAMMAD: Just water for now. [laughs]ADRIANA: Me too. Me too. I got a nice, tall glass.ANA: I think the best reminder to folks is, like, staying hydrated during the work day, and especially with the heatwave that we're having in the world right now, it's just like the best thing we can do.ADRIANA: It's true. Ana, what are you drinking today?ANA: Today we're doing...oh, you're going to like this, Guava São Paulo sparkling water from Lacroix. ADRIANA: Oh. ANA: What's your drink, Adriana?ADRIANA: Plain, old H20. Oh, and I also do have a little Perrier next to me. That's my go-to summer drink, just the lemon-flavored Perrier. I like the fizziness and the coldness. It's got the same effect as drinking an ice-cold coke but without the sugar rush.[laughter]ANA: I think the first question I would love to ask, especially knowing that you're a junior observability engineer, how did you come about the space of observability? Is this something that you had an interest in, and you'd look for some job postings? Or you kind of applied to a general job, and you ended up learning about observability and being in the space.MOHAMMAD: So the story is kind of interesting. I think we have a job posting website for McMaster, where I go to. And I came across, I think, a software engineering role for Tucows at that time. And so I went into this role, and I gave an interview for it. I think it was a technical interview. And we had a behavioral interview as well with Adriana. And at that time, she recommended me for another position as well, so observability. And at that time, I wasn't sure what it was, but it sounded interesting. So she, I think, hooked me up with some of her articles and blog posts. And I began reading into that, and I thought this would be a fantastic opportunity. So this was my segue into observability. I was really excited. And I think ultimately, I did a technical interview as well for this position, and I was able to get it.ANA: That's so awesome and rad and especially when Adriana gets to bring you on the podcast to talk about your experience. Did you have knowledge of systems prior to taking this role? Or what was your knowledge level of complex systems and running things for a company live for production and such?MOHAMMAD: So I don't know if Adriana has mentioned this before, but this is my first tech job. So this is my first actual experience in this field. So I had no idea of what it would be like to actually look into a large system. And all I had before was just schooling, so I didn't have much to go on. But it's been a huge learning curve. And I think I'm really good so far into my internship.ADRIANA: And it was trial by fire because my philosophy with my students [laughs] was basically like, hey, yo, you're a member of the team. [laughs] So you're going to be productive like everyone else. One of the things that I personally like about hiring interns is I love the attitudes of interns because it's before y'all get jaded. [laughter]ANA: Amen.ADRIANA: You work in tech long enough, and you, unfortunately, become a little bit jaded. Maybe you're lucky if you don't because you've worked at a super cool company throughout most of your career, but I got jaded pretty early. So I love [laughs] the lovely go get 'em attitude from interns. And I do love that interns are always willing to admit when they don't know something, which I think is so refreshing. Because I think a lot of people come into the workforce, especially more senior people, and they feel like, oh, well, I'm senior; therefore, I must know all the things, and then they don't ask enough questions. And, Mohammad, one of the things I appreciated when you first joined is you asked questions about everything. And it was awesome, honestly, because it's like, yo, if you don't get it, you ask a question. So then that way, we can help you do your thing so that we can help you move forward. So I love that.MOHAMMAD: Yeah, I think my mindset for this was basically to just ask as much questions as possible so I could understand more. And the more familiar I got with it, I think I wouldn't have to ask a lot of questions in the future. So my thing was just to get it over early. Try to understand everything that there is to know, and then you'll be more productive or more knowledgeable on all the material.ADRIANA: True. True. Yeah, and you've definitely gotten there.ANA: What advice would you have to anyone listening, whether they're starting in tech, getting into production systems, observability, SRE? Like, when it comes to asking questions, I personally know that I struggled feeling comfortable asking senior engineers questions when I was starting as an SRE intern, specifically where I was like, I have no knowledge of what's going on, but y'all hired me. If y'all know that I don't know anything, you will fire me. So it's hard. [laughs] How do you go about doing that? What is your advice nowadays?MOHAMMAD: I don't know if it would be advice. But I would say I was very lucky to have people like Adriana on my team who are very welcoming, and team dynamic is a huge thing. Because I don't know what I would be like if there were other people who weren't willing to help me out. But for my case, it was just a matter of luck that I landed in a really good team, and everyone was willing to help out, and they still are. ANA: That's amazing to hear, and credits to Adriana, but we're going to pretend she's not speaking right now. [laughter] We're going to have leadership folks listening to our podcasts and such, and what would you say are two things that your team did to make you feel really welcome and a safe space to ask questions?MOHAMMAD: I think, first of all, the onboarding process was really good for my time here. Adriana had a cheat sheet for me, so I was quickly able to get into some of the technologies we were using. So that was really important. And I think in my early days, I was paired up a lot with a senior engineer, so I feel like I was able to see them work. So I think that helped me to get more comfortable. And ultimately, I think now I can do work on my own. Versus when I first started and I still was getting to know everything, I kind of needed a helping hand. But I think that was really helpful in the process.ANA: Those are two amazing things, buddy systems to make you feel like a safe space to ask questions. And documentation really goes a long way. And that applies to everything in tech; documenting stuff is really key.ADRIANA: It's so true. Yeah, buddy system is huge. Because if you have someone where you feel safe to ask questions, the sky's the limit. You can be so productive, no judgment, and just do your thing. Ask your questions.ANA: I would say I've had a lot of senior engineers that don't know how to do the buddy system. Like, they don't really create a safe space. They're more of like telling you what to do, and it's not in a way of upleveling you, which I think that's quite...like, I haven't heard too much about those conversations happening of, like, how do you uplevel to be a good peer for early grads or early-career folks?ADRIANA: Huh.MOHAMMAD: I think it's just have a really nice manager like Adriana. That would be one.[laughter]ANA: Is there anything now that you've gone through some of the time of your internship that makes you really excited about being in tech or being specifically in the space of observability?MOHAMMAD: Based on the work I've been doing now and just seeing SREs and devs in our own company, I think just knowing the power of what a fully observable system can do for them because I've seen some of the hardships that they've gone through. And I know for a fact that if we had proper observability with the best standards and I think troubleshooting whatever problems they need to fix, it would be much easier than currently, it is right now. Because at this point, I think people are kind of heavily reliant on logs. And we need more people to be tracing because I feel like traces basically can tell you the same information as logs but give you more context as well. So I think that's basically what I think of this.ANA: That context [laughs] really helps when you're going through the fire when you're going through those incidents where you really can't understand why your software stopped working with the last release.ADRIANA: I think it's so cool that because you're entering into the world of observability at such a young age, you're at the perfect spot, right? Because you don't have any previous biases from the old monitoring APM dashboard is king kind of mindset, logs are king. You're coming in from the yeah, man; this is the way it's supposed to be done, okay? Which I think is awesome. And I think it's the best way to continue to foster the message is not only the paradigm shift for the people who have been doing this for a while but also the fresh, young perspective of the people who are basically, for lack of a better term, indoctrinated into the trace-first mentality of observability.MOHAMMAD: Yeah, and I'm really glad I was able to start my career off with this because I feel like in the future, for sure, this will always be a part of me to always try and include observability in whatever software that's being built or being looked at.ADRIANA: That's awesome.ANA: I think it's really fun to be able to come to a space in technology and be able to be trained and taught those best practices that can happen in a discipline. I did not know any systems, DevOps, or ops, and I came in as an SRE intern. Why would you ever put a full-stack developer as an [laughs] SRE intern? It's very interesting. But the fact that a big organization was able to trust me that I could uplevel my skills meant a lot for my career. And at the same time, it was just the ability to deep dive into so many technical concepts, and like what Adriana was saying, the fun thing with interns is just how hungry they are to learn. You are so not jaded. You might be jaded from school. ADRIANA: [laughs]ANA: But you finally get to be with folks that write documentation papers, that are working on your favorite apps, and you can be like, oh, I actually understand things now. One of my favorite things that I tell folks to do is read incident post-mortems for the purpose of learning about how things work or talking about failure in a more comfortable way.ADRIANA: Yeah, yeah, it's true. And it's so important to get that mentality early with our new grads and our interns.ANA: As you say that, I think about what advice would you have for your school when it comes to the stuff that you've been learning in your internships?MOHAMMAD: I think not just my school but a lot of schools should have observability as, like, I don't know if it's a mandatory course but at least have it as an option. Because I think, based on what I've seen this year so far, it's a huge help to at least people who might want to go into SRE and look into larger systems. And you'll need proper troubleshooting and not just dig through logs and stuff. So I think this should be a part of the curriculum in some manner. ADRIANA: That's a really cool point because regardless of where you end up in your career in tech, troubleshooting is a key aspect. And observability unlocks so much of that, and getting you into a troubleshooting mindset, not just learning how to troubleshoot but having the right information to troubleshoot, is so, so critical. And yes, a lot of it is learned on the job, but I think there's something to be said for being exposed to that even in school. Because school I found is so theoretical and so jarring when you go from like, oh, I got whatever 90% in whatever databases course. [laughs] And then you go and work with the real database, and you're like, oh, shit, [laughs] it's so much more complicated.ANA: Theory to practice is not just like a level up. We're talking about 10-15 levels of like; you can take the theory concepts and bring them into school. But when we look at most technology companies, even microservices to monoliths, there are a lot of moving pieces. Even if you work at a startup with three engineers, there's a lot of code and parts of your infrastructure that are been worked on that you ended up realizing that you may have only focused on a certain part, or you really understand compilers in code but not how anything plugs into one another. ADRIANA: Yeah, so true. ANA: For me, one of the things that I...I do a lot of work with education non-profits of bringing in folks into tech. And one of the biggest things that I am a huge proponent of is, one, internships. Getting real-life experience while you're going through education is so important to me because it's like, are you even studying something you would want to do, or are you just chasing money? Or you're like following your parents’ footsteps.Then the part around project-based learning, I think project-based learning is one of the great ways to learn together, build something, get comfortable with failure once again and iterate. A lot of folks that do those are able to uplevel the next time they get to their next class and stuff.ADRIANA: Yeah, it's so true. It reminds me of doing design projects in university, and that was the closest you got to the project-based learning where you're trying to have a real-life example of stuff which is cool, very stressful too. [laughs] One thing that I wanted to ask so, Ana, you know, being in the States, is co-op a thing in the States? I wasn't sure. It's a huge thing in these parts and in Toronto, definitely Toronto area. Is it a thing there? Or is it mostly just internships between, like, in the summer kind of thing?ANA: You're making me go back to when I first heard about University of Waterloo and learning about their co-op program and being angry that I wasn't going to a co-op school. It was one of those things where it's like; I would help hire interns at Uber. And we would look at their resumes, and it's like they're coming in with four internships at companies, and a lot of them were big tech, of course. And you just spent a year working in the field but working with so many amazing minds.Like, of course, I would love to just interview you and get to know of you…

    Full show notes at the publisher

    Cloud Love with Renata Rocha Oct 25, 2022
    Show notes

    About the guest:Originally from Rio de Janeiro, Brazil, Renata now calls Toronto, Canada home. Fun fact - she has been working in the tech industry since the last century! Renata started her tech career doing sysadmin work for ISPs in the late 90s. In 2010, she got hooked on Cloud when a co-worker told her about "this new Cloud stuff". She hasn't looked back since.Renata loves being hands-on, taking any opportunity to get her hands dirty on cool new tech to experience it for herself. She also enjoys working with high-level enterprise customers and on the business side of things, namely understanding constraints and limitations, solving the puzzles that those scenarios represent, and finding smart solutions that make everyone happy.In her spare time, Renata likes to run, ride her bike, and watch a ridiculous amount of movies. She also loves cats.Find our guest on:Renata’s TwitterRenata’s LinkedInFind us on:On Call Me Maybe Podcast TwitterOn Call Me Maybe Podcast LinkedIn PageAdriana’s TwitterAdriana’s LinkedInAdriana’s InstagramAna’s TwitterAna’s LinkedInAna's InstagramShow Links:Amazon Web Services (AWS)CloudSREInfrastructure-as-CodeGitOpsKubernetesOracle CloudAmazon SageMakerVirtual Private Cloud (VPC)Alibaba CloudTranscript:ADRIANA: Hey, everyone. Welcome to On-Call Me Maybe. This is the podcast where we talk about DevOps, SRE, observability principles, on-call, and anything in between. Today we are talking to Renata Rocha. She is a DevOps Manager. And we are here to get her insights on all things DevOps and tech. I am your host, Adriana Villela, with...ANA: This is Ana Margarita Medina, also joining in today.ADRIANA: And we have a very Latina contingent today, very exciting. All girls, all Latinos. So woo-hoo, represent, y'all. Renata, why don't you tell us a little bit about your work, how you got into tech. What's your story? RENATA: Sure. Hi. Thanks, everyone, for joining us today. Like Adriana said, I'm Renata. I'm a DevOps manager, and I have been working in tech since last century. And this is quite an interesting story because we are both from Rio, so Adriana probably knows this much better than everyone else. But Rio de Janeiro, back in the day, was not the greatest place for technology. And I worked in very interesting places like ISPs as a sales ops; that's how we called it back then. And that's how I got in touch with Linux, Unix. And then, I came to Canada about 12 years ago. And I had a co-worker who told me about this cloud thing, his words, not mine, okay? And I was like, "Okay, show me this cloud thing." He showed me AWS, very early days of AWS. And I was like, whoa, mind blown, okay. This was like, wow, how can this be a thing? And I started playing around running some instances very early days of AWS. Keep in mind it was very, very small, not like 1,000 services like we have today. I was absolutely fascinated; no coming back from that. You can imagine once you see it, there's no way to go back to what you were doing before. And that's how I got into DevOps now, new word, new world for me, not only a new word but a new world for technology. And I started doing it by early 2010. And I have been doing cloud engineering, DevOps, SRE for 12 years now since I came to Canada. And I love everything about it. I feel like there are endless possibilities. And I also love how fast-paced it is because every day seems like we have a new cool technology showing up, and I have the opportunity to learn it because everything is so open. Everything is available online for me to try, for me to learn. And I love learning. I love challenges.ANA: That's an amazing introduction. I always love hearing how folks got started in cloud. And it's definitely one of those things that it's this cloud thing. Like, until you decide to take a look and you start understanding, it's like, oh, you're just calling it a data center, a thing on the cloud, got it. Now we can actually still use the same resources, but it takes a lot of less ops work. [laughs]RENATA: Exactly. It makes things so much easier. That's one thing that I particularly love about this cloud thing. Let's keep the support because I love it, how encompassing the cloud thing can be. Because people think that to work in technology, they need a degree; they need to spend lots of money to go to school when studying computer science, engineering, something like that. And that cloud thing, there is no school that teaches it. It's the school of everyday learning. You have to go online, and you have to learn for yourself. You have to find the things that you like and just dedicate yourself to that. None of the things that I use every day at my job I learned at school, the university. I did mathematics. Nothing that I use at work was taught to me at university. So that is the amazing thing about cloud engineering. You learn by yourself every day, and it's an ongoing learning thing. So I find that it embraces this culture of open knowledge, and it's very diverse. The workplace can be very diverse and can be very welcoming to all people of all backgrounds.ADRIANA: I love the fact that where we started our careers, and I think it's probably true for all three of us, where we started our careers is not at all where we have ended up. I mean, this stuff didn't exist in school. I don't even know if they teach cloud stuff in school, to be honest. And it's okay because it's one of those things that you pick up. The other thing I wanted to mention, and I'd love to hear your thoughts on this, like, I was so intimidated when I first heard about the cloud. And [laughs] I started...my first cloud was GCP. And I'm like, oh, that's it? Really? [laughter] Just some cool commands to bring up some resources that I don't see but exist out there somewhere.RENATA: That's a very interesting thought. Because when my co-worker told me about it, it took me a while to wrap my head around the abstraction of what the cloud was, especially having so many years of experience with on-premises only. I couldn't really understand how such a thing was possible, feasible and above all things, have a vision of how that thing was not going to be only a fad.And my co-worker told me, "No, no, no, this is the real deal. This is where we are going in the future." And he insisted with me, "No, no, you have to learn this. Keep your eyes open." I'm very thankful that he insisted that I learn that because that was going to be the future. And I was like, okay, I have to regroup my thoughts here and go back home, create an account on this thing. And I was like, Amazon? But aren't they a store? Nothing made sense. And this is an interesting tale for people listening to this who are new and who feel a little bit confused about cloud. So it's normal. If you feel confused, if you feel like this is wrong, that this is not for you, it's normal. I felt like this for the first time, and here I am today. So it's a journey, man. It doesn't happen from one day to the other. You don't start today doing this, and then the other day, you're going to be the master of the cloud. And then I was like, oh my God, it's not like my Amazon account; it's another thing. But it's also Amazon, so many thoughts. Everything is happening at the same time. And then, I created the account, and I started clicking buttons. I'm like, so I don't write a code for this; I click buttons, whoa, okay. It was very confusing. It was very crazy. But like I said, when I finally understood, it was like, whoa, my God, I understood what he was saying about this is going to be the big thing because it was so much simpler, and it was very powerful. And yeah, of course, I can see how easy it is and what I said before about endless possibilities. Once you do one, and then you do it again, and you see yeah, okay, I don't have any limitations on how...because you are a sysadmin. You need to add more memory to a server. And it's a whole process, and you have to go to the data center. You have to buy more RAM, and then you have to stop the server to add more RAM. And to do that on the cloud, it's a click of a button. Or you just spin up another instance, and you group things. Like, of course, it just makes sense. And that's how I saw that was going to change my life. And that's why I try to tell kids who are in school...to your point, I don't think schools are teaching that because I interview a lot of kids who are just fresh out of university for computer engineering or for science and their understanding of cloud engineering, SRE, DevOps is very, very bare bones. They are very good at programming. They're very good at algorithms, AI. But cloud engineering is still very, very lacking. I don't know how it is in the U.S., but Toronto universities they're not there yet, which is a pity. So many companies here in Toronto have bridging programs. They teach a bootcamp for students that are just getting out of university. So they hire these students. It's not like an internship, so the students get paid to work, but they learn for, let's say, a couple of months. I'm not going to name names, but I know big companies that are doing this. So you do a bootcamp for like six weeks, and you learn things. You learn how to work with cloud engineering, and then you are bound by a contract that you have to work for this company for about a year or so. But a number of companies are doing this because the universities are not teaching the students the skills that are required in our market right now.ANA: I actually really liked that flow. I'm assuming this is different than the co-op flow that a lot of the Canada universities have, like, the bootcamp program. That's actually really awesome because as we're talking about this, I also have the same thoughts like; coming into Cloud, DevOps, SRE, extremely hard unless you have the right people in the room to ask questions and to be in a place that's inclusive, and you're not going to get dinged for asking, quote, unquote, "questions that make no sense," aka stupid questions. [laughs] It's very much of like, I think, when we think about the DevOps movement, it is very people-oriented. SRE is very much around empathy and working with others, that when I think about actually putting this in a class course, it's just going to be theory once again. And there's a lot of stuff in cloud engineering that doesn't make sense unless you start running environments when you're working in dev with 5 to 100 engineers. And when it comes to SRE, if you ain't got customers, [laughs] does reliability really matter? RENATA: Big-time customers, right? One engineer in my team had a lot of theory experience, and he had great grades that he had in his resume. And I was like, "Yeah, cool stuff, bro. Let me just tell you something; it is okay if you feel your first project is not the same as your school experience. It's okay for you to tell me that you're struggling because it happens to everyone." And he said, "No, no, I was a great student in school. I had the best grades. My GPA was great." "But feel comfortable to come to me if things are hard." Of course, in the third week of the project, he came back to me and said, "Hey, remember what you told me? Things are hard. I'm so lost. I don't know what I'm doing." And I said, "Yeah, thank you for coming to me because let's try to debrief what is hard, what is difficult here. So let's have a one-on-one call here and try to understand what's going on. And thank you for sharing with me that you were struggling." So yeah, and I think this is what is great about being an engineering manager. You can share your past experiences because it's happened to me. And now I'm talking to the younger generation and seeing a little bit of myself in younger folks.ADRIANA: That's so cool. You know what? I love the point that you made about the student saying, "Hey, I was a great student, got great grades, so this shouldn't be a problem for me," because it makes me think of when I came out of school. And I soon came to the realization that I was excellent academically but made a terrible, terrible worker. And it wasn't until I came to terms with that that I was able to really grow into myself as a person in technology. And I think Ana and I have talked about psychological safety before. And I love the fact that you offered help to this person who worked for you, saying, "Ask me questions." You gave them a safe space where he could come to you for help. That's another thing, a lot of people coming out of school they have this impression that we, as more senior people, expect them to know all the things and therefore, they can't ask questions, or they can't look dumb. And we have to make it clear. We have to foster this culture of, like, no; it's okay to ask questions. It's okay to not know everything. And if you don't ask questions, then how the hell are you going to do your job effectively?RENATA: Yeah, exactly. What I try to do as a manager...like, I was not born a manager. One becomes a manager from learning from your own mistakes. And what I try to do as an engineering manager is to foster an environment that is safe for people to ask questions. And I tell people that there are no wrong questions. The only thing that I don't like is for people to not ask questions. Ask questions always, and it's okay to not know the answer. If I don't know the answer, I'll say, "You know what? I don't know the answer for that. But I will try to find the answer for you." Or maybe we can try to find the answer together. And I will try to redirect the person to someone that I think knows better than I do because I'm not a person who knows everything. And I'm not an expert in all subjects, which is also okay. I know a lot about certain things, but I don't know everything about everything. It's also okay to make mistakes, and making mistakes is how we learn stuff. Falling and getting up again is how you get better at things. One thing that I like to tell people is that I have danced all my life. I did ballet all my life. Ballet is beautiful. But the thing about ballet is that it's incredibly hard, and there is no such thing as a perfect ballerina. If you do ballet a lot, you can point at the mistakes that people were making in every single video that you watch. And ballet is a journey. It is a journey of trying to perfect yourself. Very basic example is watching the video of the Swan Lake. And the thing with that is that the ballerina is trying to mimic the movements of a swan dancing; the Swan Arms is like the classic movement. But, of course, you are not a swan. You are a dancer doing Swan Arms. That is incredibly hard. And you are not able to do the same movement over and over again. And you have, like, for the ballet, doing those movements. And that is a journey. You are not going to be able to do that every single time the same way again. And the thing that you have to see is that you have to enjoy doing that. And you have to learn every single time that you do that movement over and over and over again. It's a journey. It's enjoying the journey. It's learning with every mistake that you do. And it's the pleasure that you have with learning things. Okay, if you don't enjoy learning, then why are you doing this? And you have to ask questions because if you don't ask, "What am I doing wrong here?"…

    Full show notes at the publisher

    Improving Quality with Observability with Parveen Khan of Thoughtworks Oct 18, 2022
    Show notes

    About the guest:Parveen is a UK-based senior quality analyst consultant at Thoughtworks. Being a quality advocate, she believes delivering high-quality products is everyone's responsibility. She loves collaborating with teams and optimizing processes, tools and methodologies to enable the creation of high-quality products. She is also an international speaker sharing her stories and experiences in testing to inspire other people around the globe. In her spare time, she plays the role of wonder woman for her two lovely kids. Find our guest on:Parveen's TwitterParveen's LinkedInParveen's BlogFind us on:On Call Me Maybe Podcast TwitterOn Call Me Maybe Podcast LinkedIn PageAdriana’s TwitterAdriana’s LinkedInAdriana’s InstagramAna’s TwitterAna’s LinkedInAna's InstagramShow Links:ThoughtworksExploratory TestingAdditional Links:Blog Post: Observability for TestersBlog Post: Why Observability Matters for TestersO11ycast Podcast: Ep. 26, Unknown Unknowns with Parveen Khan of Square Marble TechnologyParveen’s Speaking EngagementsTranscript:ADRIANA: Hey, everyone. Welcome to On-Call Me Maybe. I am your host, Adriana Villela. And with me, I have Ana. I'll let Ana introduce herself.ANA: Hi, y'all. My name is Ana Margarita Medina, and I'm very excited. We're going to have an amazing episode today. So I'll switch it on back to our co-host, Adriana.ADRIANA: Today with us, we have Parveen. I'll let Parveen introduce herself. Why don't you tell us a little bit about yourself and also how we connected, how we found each other.PARVEEN: Yeah. Hello. So I'm Parveen Khan, and I'm a senior QA consultant at Thoughtworks. Yeah, I kind of work as a quality advocate, all about quality with the teams. And also, I share my learning experiences across trying to speak at different conferences and writing blog posts. And I share some of my thoughts being on a podcast. I'm super excited to be here today. So I still remember such a simple conversation, like, I think I read a blog post about observability myths. So that's where I think I thought, okay, this is so cool. It's such an easy way to explain all this all-around observability. Because I know there are a lot of different blog posts available, a lot of resources there. But then I think when I read that blog post, I was so fascinated that I think I reached out to Adriana just on LinkedIn, saying that, "Oh, this is so cool. You've given such great information out there in such a simple way so that everyone can understand." So I think that's how we kind of met.ADRIANA: Yeah, totally. And I love that we were able to connect through LinkedIn. And it's funny because you're not the first person that I've connected with through LinkedIn because of the blog posts that I've written. So I really appreciate it when people reach out to want to chat about these types of things. And I'm super stoked that the blog posts on observability resonated with you. And after you reached out, we booked a meeting to talk about observability and QA. And you sent me on this whole, like; you opened my mind to this whole new way of looking at observability and QA. Maybe why don't you talk a little bit about how you started applying observability as part of your role?PARVEEN: For me, I think it was never like, okay, this is what observability is, and this is how you can use it, and this is how I can try to make use of it as a QA; it was never like that. So I think it was a way of me trying to learn different challenges at work. I remember I was working with one of the teams. And I came across a lot of challenges, which I was unsure of why these things are happening. We kind of had no information at all. It was interesting, like, I was just trying to read a lot about observability at that point of time.I just came across this new term observability and then how it kind of connected between...sometimes it's like you cannot understand when to use it until you are in that situation. So I think for me, it was more of like I was in that situation where I realized that, okay, this is the reason why we need observability. So I think that's where my learnings and my understanding started from, and then that's where I started to learn a little bit more about it. Because I think for me, initially, it was more of like, okay, this has nothing to do with me. I don't know if this is something as a QA I can make use of it, or I can try something on it. But I think the more I tried to learn and explore more about it and found those challenges on the product that I was working on; I think that's where it opened up a lot of possibilities for me, okay, I think this is how I can use this being a QA, and this is how I can try to add some value from a QA perspective. So I think it started in that sense for me.ADRIANA: That's so cool. And I think it's interesting, too, when you talk about trying to understand what observability is because I don't know about you, and maybe Ana, maybe you can relate to this as well. Like, when I first heard the term observability, and it was, I would say at least like two years ago, I swear to God it took me so long to wrap my head around what it was. I don't get it; I don't get it. I feel like it's something that's important, something really cool that we need, but I don't get it. And all the definitions were so academic, and I'm like, ah, what the hell is this thing?ANA: I can relate 10,000% about everything that was just said in the last few minutes. Because as Parveen was saying, that part about learning you hear a word and you're just like, what is this? This sounds interesting. Let me go research and see how this ties into the world as I know it and what you kind of call a mental model. And then you start asking more questions, and you're like, oh, idea, now this all starts clicking. And that's what a lot of DevOps means to me, in my opinion.You hear the word DevOps; you have to kind of come together, and it's like, dev, ops, collaboration, communication, tada, we have amazing things. And with observability, I felt like that's a lot of what the space has brought. For me, I have that similar experience where I heard the word, and I was like, this means nothing to me; move the page, keep on going. I come from chaos engineering, like; that was my introduction to systems, chaos engineering infrastructure. And I was working at Uber during that time, and we had Jaeger. So I started learning about tracing without knowing anything about observability. And I was like, we had our internal tool for metrics, and we had M3 and Jaeger. And then all of a sudden, I was like, I like dashboards. These things [laughs] make a lot more sense. I come from a perspective where I was like, I understand observability to the point that I need to, but I don't need to dive into it. And then I'm coming back to the table four or five years wiser. And it's like, oh, actually, the more we know of our system, the better we're able to understand the system that it engages with on a day-to-day basis of our users. And what happens internally is when we have our very complex architectures of like system 1 calls database number 20 but then goes back to database 1, what happened in this time span of five seconds? [laughs]ADRIANA: Yeah, totally, totally. And for me, it was like this, oh, I get a holistic view of my system? Because from my personal experience, I remember back in the day, even seven years ago, I remember I was helping troubleshoot an application. It was a vendor application that we had running in prod, and it was slow. I talked to the database person, and they're like, "No, the database is performing fine." I talked to the network person, "No, network is performing fine, hard drive's performing fine." I'm like, "Guys, something's not working, [laughs] but everyone says that their thing is running fine. What the hell is going on?"And I feel like observability unlocks this new level, expert level where all of a sudden you're like, oh my God, finally. I can find that little needle in the haystack. I have that visibility into the thing that's not working for me properly. Parveen, when you and I were chatting initially, that was part of the aha moment for you. As a QA tester, you're like, oh, I know what's wrong.PARVEEN: [laughs] Yeah, exactly. I think it's about figuring out how to know where things are wrong. It gives you the ability to look for some more information, look for trying like, okay, it's not about just saying...so as a QA, when you just go to the developer and say, "Oh, something is broken," it doesn't give you anything. Okay, something is broken, but what exactly is broken? So I think as a QA, I'll be in a better position to give a bit more details around what are the things happening behind like, you know, what has broken. So it's not like me giving the solution of oh, here's the thing that is broken, but trying to give more information around okay, I've observed these kinds of things, or I have noticed these things when I was trying to see where something has been broken. So I think it's about trying to help navigate or trying to give a bit more information to the developers.ADRIANA: Yeah, totally. When you and I chatted...I had a role as a QA tester early in my career. And I remember when I was testing, I'm like, it's broken, but I don't know why. And I'm like, I don't want to just sit here until the developers figure out what the hell's going on to find out what the problem is. I want in. I want to figure out if there's something I can do to find the problems. So I wish that observability had been as mature as it is now 20 years ago when I started my career. That would have been so awesome. [laughs]ANA: I have a similar experience of being put in a QA position where just like, here are the screenshots, here are the steps I took. Here's what didn't happen. Oh, you won't get back to me for another week? Yeah, let me try to recap all my notes when you do get back to me on why it was pink and not blue. I don't know [laughs] what to tell you; the system just didn't handle it properly. So I think the more that we're able to know the right information at the right moment, it's also critical. It is not just about knowing this information. Because when we talk about having to fix something when it really matters, like being in those moments of your customer is reeling bad, how do you make sure that we can get them the most assistance? Or we're really close to launch; how do we make sure that we finish the backlog of your tickets? So yeah, that moment where it's like, oh, it's crunch time. Why didn't this work? Why is Bobby mad at me right now? And why is my supervisor, Veronica, still questioning, like, "Hey, why don't we have this case closed in so long?" It's really getting that context when you really need it.ADRIANA: Yeah, I totally agree. And it doesn't apply just to the pre-prod QA, either. I mean, we're basically testing when we're in prod every day because our users are always finding new and interesting ways of using our systems in ways that, as a developer, you're like, oh, I didn't even think that that could be done that way, so you get the insights both in the pre-prod and the post-prod, I guess, areas for your system. And that makes it very valuable on both ends. And I think having the additional insights in the pre-prod stages won't make a perfect product, but it will certainly help to get rid of some of the wrinkles before you go into production, which I think is awesome.PARVEEN: Yeah, it's more of like, you can learn from the production system because that's where the actual things happen. That's where our users are using our features. I think it helps us in trying to understand not just how our systems are behaving in the production system but also how our users are using our production system. And then using that as feedback and trying to add that into our process. And then I like to say this, like; you cannot shift right until you shift left. You have to shift your mindset to the left and try to think. That's why I like to say that observability is more like a mindset change. And it's more of like; you try to think whenever you're releasing the feature to production, it's not about keeping your fingers crossed and saying that "Oh, I hope everything goes well." It's about asking those questions beforehand and saying, "Okay, if I release this feature to production, how would I know something has gone wrong?" Again, I'm not trying to say that observability is rocket science, and you will know everything with it, but still, you are prepared. At least your aim is you don't want to fail. But your aim would be something like, even if you fail, do you know how to get some information about it instead of blindly going around and trying to see, like, I don't know where to go kind of situation, right? ADRIANA: Yeah, totally. It's the idea of it's enabling you to fail fast, right? PARVEEN: Yeah.ANA: Beautifully said. Like, fail fast, get comfortable with failure because if we're not comfortable with failure, you're going to end up having more failures in the moments that matter to your customer or any of your users. And it's a lot of what I've been iterating for the last four years but from a whole different angle of like, the faster your engineering team is able to get comfortable with failure, the easier it's going to be when your pager goes off when that incident gets started because your team already knows what to do. They know, oh, this is how my observability tool works. And hopefully, you're working in an organization that only has one, not like five that you have to be like, [vocalization] was it this one? Or did we migrate this service to our new one? And it's those little things that really add those extra minutes, those extra hours to an incident getting closed or any type of ticket being closed and a customer being happy once again and staying as a return customer with oh, what I came to shop for actually got into the car. I got a tracking code, perfect. Or where is it that it's failing, and how can we make it be better? And as Parveen was saying, very much of injecting it and, like, knowing what's going on the closer on the shift left is the most important. I'm a huge fan of having perfect, ideal world DevOps in an SRE world across all your cycles, but I know that's extremely expensive. So it's like having it in pre-prod and staging; the more context and observability that you have, the more comfortable you are with failure. You're going to start seeing that your team is going to have a lot less of those really long outages when you are in production. And as you see organizations start doing the work, you're just like, I'm cheering for them on the sidelines. Or when you hear about a really expensive outage, you're like; we can do better. We got this. And now you're just trying to herd like 500 engineers at a big org, like; I'm cheering for y'all. Like, #hugops, you got this.ADRIANA: It's so true. I think that goes with the mindset shift that Parveen was talking about earlier, where making it a safe space for people to feel. I think observability helps provide that safe space. But then I also think that it's up to leadership to allow for that safe space. Because I've been in organizations where I…

    Full show notes at the publisher

    Finding Humanity in Incidents with Nora Jones of Jeli.io Oct 10, 2022
    Show notes

    About our guest:Nora Jones is the founder and CEO of Jeli.io. She is a dedicated and driven technology leader and software engineer with a passion for the intersection between how people and software work in practice in distributed systems. She created and founded the learningfromincidents.io movement to develop open-source cross-organization learnings and analysis from reliability incidents across various organizations and the business impacts of doing so.Find our guest on:Nora’s TwitterNora’s LinkedInFind us on:On Call Me Maybe Podcast TwitterOn Call Me Maybe Podcast LinkedIn PageAdriana’s TwitterAdriana’s LinkedInAdriana’s InstagramAna’s TwitterAna’s LinkedInAna's InstagramShow Links:JeliThe Howie GuideLearning from IncidentsTwitter thread summary of our chat with NoraTranscript:ANA: Hey, y'all. Welcome to On-Call Me Maybe, the podcasts about DevOps, SRE, observability principles, on-call, and just about everything in between. Today we're talking to Nora Jones, CEO and Founder of Jeli.io. And we're very excited to have you here. Thanks for joining us.NORA: Thanks for having me.ANA: We have the first question for today's show; what is going to be your drink of your choice today?NORA: I have a smoothie that I'm working on. So it's from the shop across the street. It is like a little kind of milkshake smoothie. It has some oat milk, some strawberries, some bananas in it, and it looked delicious. So they sold me, and I now bought a $12 smoothie that I probably could have made for $3. [laughs]ADRIANA: Sounds very delicious.ANA: What are you having, Adriana?ADRIANA: I am having some good, old water.ANA: Good hydration. Similarly, I'm doing sparkling water with a little raspberry cranberry flavor, just trying to beat the end of summer heat.ADRIANA: Yeah, it's pretty hot where I'm at. We're like high 80s, [laughs] so it was a very hot day.ANA: So, for our listeners today who are not familiar with Jeli, can you tell us a little bit about the company you have created?NORA: Jeli is an incident analysis company. First and foremost, we cover really the whole incident management suite. But we go deep on an area that I think has frequently been not very much talked about in tech but is incredibly important, and that's on understanding what happened after we had an incident. There is so much learning to unpack with how we coordinated, how we spoke to each other, who we brought in, why we needed to bring them in, what systems were involved that doesn't really get talked about. In my experience, the tech industry really focuses on what we know best, which is the technology and the software behind what led to an incident. And there's a whole other story that didn't get unpacked, which is how the organization works together, the relationships, the psychology aspects, which I think can sometimes feel really awkward to talk about, but it's also really important to talk about. And so that is really what we focus on is making that easy for folks to unpack in a way that can help them improve in the future.ADRIANA: That's so cool. So you actually tackle the sociotechnical aspect of it, too, then?NORA: Yeah, absolutely. That's our main focus, and it's a lot of fun. But we're also all engineers. We've been a part of the big technical aspects of it all. And we've seen the social aspects not really be spoken about very much. I think it's a big miss. And it's honestly a business advantage to be able to talk about the social aspects as well. So we're really trying to give every company that advantage.ANA: Especially when it comes...like, a tool is not even necessarily just a new culture that you're trying to bring into organizations. You're kind of bringing in the new culture, but it's being facilitated with something that's going to help them.NORA: Totally. It's definitely meant to help them. It's meant to take typically sterile and maybe sometimes boring or a very emotional thing that happened, which is an incident, and make it fun, make it a learning opportunity, make it not feel like a bad thing that happened but something that is expected as part of moving fast in a very technologically advanced world, and just really trying to extract value out of it and also celebrate your employees and your organization along the way.ANA: As you've been working on this field, do you feel like the definition that we have for incident management, incident response needs a refresh when we look at it since we have been focusing very much on technology broke; it didn't matter how it broke?NORA: Yeah, 100%. I think it needs a huge refresh. I think there are a lot of things that need to change; I think around how we view incidents, I think around how we speak to people that participate in them afterwards. I think it's a lot of emotional burden to carry fixing an entire incident when you are also an expert in the situation.But it's like, it also needs to be approached with care when you talk with them afterwards. And I don't feel like that is always spoken about as much, but it's really, really beneficial to do that. So I think there are a lot of things. I mean, at Jeli, too, there are even subtle language changes we make in the tool. Instead of using the word incidents, we use the word opportunities. So like in CRM tools when you sign in, it shows you your opportunities rather than these are the people you need to sell to and meet up with. It's like, that almost feels kind of not fun and kind of a little transactional whereas opportunities it's like you're building a relationship. You're building an opportunity to grow. It's similar with incidents; it's an opportunity. It is an opportunity to grow. It's not a bad thing that happened that you need a checklist to recover from. It is like a way to evolve. It sounds like such a minuscule language shift. But doing those little language shifts can actually really help people think about them differently, which will help the org get more out of them at the same time.ADRIANA: I love that because it turns it from a glass half empty scenario to a glass half full scenario, right?NORA: Bingo. Yeah, yeah, exactly. It's way more like this happened. It wasn't fun. But we can collaborate on it, and we can just treat it as an opportunity to learn and grow. And it's something I've thought about for a really long time, and something I used to get hired into organizations to help was to change the culture and the way we spoke about incidents, even in one on one conversations, like passing by, getting coffee together.But I really had to walk the talk in my own organization too. I'm making a tool that I'm trying to sell to people [laughs] to help them think like this. I'm going to have to implement it internally too. And I will earnestly say incidents are kind of fun here. They're never expected, but they are... we really enjoy working together and collaborating with them. And they're just considered a regular part of work and not a very big deal. There are certain incidents, of course, that you cannot avoid in almost every company, in every industry that is successful, that are going to be painful and are going to be a big deal. And I think the way you talk about them afterwards can still help folks feel psychologically safe, help your retention, help challenge your engineers and all the people in your organization to build their expertise in a way that helps your organization long term too.ADRIANA: That is so cool because I think it taps into what Ana was saying in a previous blog post about basically leading with empathy. And this really taps into leading with empathy. And I love, too, that it sounds like it takes the PTSD out of dealing with incidents because of the positivity around it, which I think is so cool. Because I think so many of us have been so burned and so scared to even own up to mistakes because of repercussions, right? Oh my God, I caused this. Now I'm super screwed, like, fingers pointing everywhere, blah.NORA: It's almost worse when you're feeling that way, and then someone else gets assigned the post-incident review. And you were just waiting to figure out what they're going to say about your enrolment and involvement in the situation, even though you did what made sense at the time. You care about your job very much; I'm sure they do as well. And you just have this anxiety, and usually, it's about how they're going to fill out a pre-templatized Google Doc. And oftentimes, it ends up with a lot of their opinions of the situation. And so you're almost like subtly DMing them, trying to schedule meetings with them to try to make sure you're involved in the conversation because you very much should be. And the thing is we make all that easier, like, we make that easier for the you in that situation. We make that easier for the writer in that situation. We make that easier for the person that created that templatized process. And I think a lot of the problem is how we've been set up as an industry, too, which is, yeah, use a Google Doc and Slack, two tools that were not made at all for incidents, [laughs] and come up with the answer to what happened in this incident, which doesn't make sense. It doesn't set anyone up for success. So what we really do is we have a couple of different personas that we really tailor towards. We are there for the person that is creating the post-incident review and is tasked with this, but we also assume that they have a lot of other stuff going on. And they may not get your full account in the way that you deserve and the way that the person deserves. And so we really help them see the key players in the situation and the key standout moments so that they can collect all the perspectives and get all the different views of expertise because ultimately, incident review shouldn't be done in a vacuum by one person. They should be a collected, amalgamated experience that is put together. And the thing is, everyone's individual experience is incorrect, and it's also very correct. And so the role of the incident reviewer is to collect all those and highlight the differences between those experiences because the real answers are in the deltas, like how you viewed what happened versus how I viewed what happened. And again, neither of us are wrong, and neither of us are right. And it's very hard to do that as an incident reviewer, but that is what we try to make really easy. So that, what you said, it feels psychologically safe afterwards, that there's not a lot of emotional burden afterwards. And that you really feel like it's kind of a team effort that you're working through together.ANA: The post-incident reviewer is actually able to see you in a sense like in a human aspect.NORA: Right, exactly.ADRIANA: How do your customers respond the first time that they use that approach to handle these opportunities, as you put them? What goes through their minds?NORA: I would say our first initial users, like all of our initial inbound customers, were folks that were very bought into this way of thinking, but they might have been one of two people in their org that was bought into this way of thinking. And I used to also be that person in this org. So I really tried to make a tool that would help them socialize it better and show the ROI of it better. And so I think the initial reactions, especially from those people, are like, oh my gosh, now I can screenshot this thing where you're visualizing this for me, and you're putting this together in a way that multiple different parties can understand. I feel like the initial reaction has been really great. And I think a lot of folks in the industry, when they hear the word learning or learning from incidents, sometimes the initial reaction is like, but what does that mean? Like, what does learning mean? What is the ROI of that? And we really try to show the ROI of that in a way that can make it easy for those skeptics to go, oh, interesting. And so our goal is to help the person that already really is trying to move some of these processes and thoughts forward into their organization and really help them socialize it to their colleagues in a way they understand how it impacts their work positively too. So it's like there's an initial amazing reaction towards the people we know and interact with, and there's a slow trickle through the rest of the organization that is really quite cool to see.ANA: For anyone that's trying to convince their manager to start having this culture change of seeing them as learning opportunities, but their manager is kind of pushing back because, like, the same failure doesn't happen twice; we can't learn from it. What is your take on that?NORA: My take has honestly evolved a bit over time. I was very much in the like; I will stand on this hill and die on this hill about some of these things. And now it's like, that's not the way to make change in your organization. I think everyone has...are y'all familiar with the terms blunt end versus sharp end in terms of expertise?ANA: I think our listeners might not be. We'll definitely get the explanation. NORA: Okay, yeah. So just to distill it just in a quick sense...and there are much better infographics online about it. Richard Cook has a really great one where he talks about Above the Line, Below the Line, and I can send a link to it afterwards. But the sharp end is like when you are in a role, and you are seeing all the depths of your role. And then the blunt end is more like what others might see about your role in the organization. I think the thing that's evolved for me is everyone has their own sharp end beyond engineers. And I know that sounds very obvious, but I think directors and managers, even the curmudgeon-y ones in those situations they, also have their own sharp end. And I think just getting curious about everyone’s sharp end in an incident is what entices change. Some initial mentalities in this space have been like, let's completely just only listen to the engineers in this situation. We really have to understand their expertise, and that is totally true. And I think sometimes managers and other folks’ sharp ends we're left out of that conversation, which does not make change possible. And so it's like getting curious about those sharp ends as well so that they can all integrate together.ADRIANA: And I think that plays nicely into what you were saying then about, like, everyone's perspective matters, right? NORA: Yeah, it matters. Yeah, it's like, I don't want to say wrong or right, but everyone's perspective is incomplete, and it's incomplete in their own way. And so they all need to form together to be complete. And no one's is more right or more important. It's just kind of collecting this together, which requires a really skilled facilitator, which that's what a lot of companies don't invest in. They're like, "So, and so you were the incident commander here, so you get to do the post-mortem. And you have eight hours to do it for our eight-day incident. Yeah, let us know all the stuff we need to do afterwards." And then they'll be working to complete it really quickly. They won't get everyone's perspective. They’ll hoard a quick post-mortem. Then they're like, "Yeah, so and so that also did this; you also have to do all these other things this week that have nothing to do with this incident. But we really want all the action items from this incident too." So they quickly rushed to do those. And then those action items ha…

    Full show notes at the publisher

    OpenTelemetry & Nomad with Luiz Aoqui of HashiCorp Oct 04, 2022
    Show notes

    About our guest:Luiz is a Toronto-based senior software engineer at HashiCorp working with distributed systems on the Nomad workload orchestrator. Before that, Luiz was a full stack and DevOps engineer at IBM, leading a team that builds and manages a SaaS e-learning platform. Find our guest on:Luiz's LinkedInLuiz's TwitterLuiz’s GitHubFind us on:On Call Me Maybe Podcast TwitterOn Call Me Maybe Podcast LinkedIn PageAdriana’s TwitterAdriana’s LinkedInAdriana’s InstagramTed’s TwitterTed’s LinkedInShow Links:IBMHashiCorpNomadWorkload OrchestratorKubernetesRancherPivotal SoftwareRedHat OpenShiftPodmanOpenTelemetryQUEMUOpenTracingDistributed tracingTranscript:ADRIANA: Welcome to On-Call Me Maybe. I am your host, Adriana Villela, joined by...TED: @tedsuo on the internet, Ted in real life.ADRIANA: Awesome. Today we have...LUIZ: Hi. I'm Luiz.ADRIANA: So, Luiz, tell us a little bit about yourself.LUIZ: Sure. My name is Luiz. I'm an engineer at the HashiCorp working in a project called Nomad, which is our orchestrator solution. Yeah, I've been working on the project for almost three years now. And before that, I was in that developer operations space, meaning that my team was not large enough to have an ops team, so all developers had to do a little bit of [chuckles] operations and everything. So yeah, that's how I got involved in this space. And then eventually, I got the opportunity to work at HashiCorp and its tools.ADRIANA: Cool. That's awesome. And it's funny how you and I met because I think we met on Twitter [chuckles]; if I'm not mistaken, though, I think my post on HashiCorp my explorations of Nomad...because last year I was a total Nomad noob. I was at Tucows last year running a HashiCorp team, all things Hashicorp just about. So I was like, oh, shoot, I better learn how this stuff works. I guess that's how we met. Now we follow each other on Twitter, which is awesome. And I guess we also have...there's the additional HashiCorp connection because I think, Ted, you said you worked at HashiCorp at some point, right?TED: Ah, I did not actually work at HashiCorp. But when I was interviewing, when I was looking for my last job, it came down to either HashiCorp or Lightstep. Both were really interesting to me. I like the idea of bootstrapping up the OpenTracing project, which is why I went with Lightstep. But I've always enjoyed HashiCorp's approach to engineering and product development. And Nomad actually was the project I was most interested in because, in my last job, I was working on container scheduling at Pivotal on a project called Cloud Foundry. So I really enjoyed the domain space of scheduling and building that part of a distributed operating system. I thought it was really cool.ADRIANA: Yeah, and I have to say, coming from a Kubernetes background and being thrust into Nomad, I was like, oh, man, this is so much easier to get started. [laughs] My mind was blown right away. It was awesome. It was awesome. I can understand why HashiCorp has such a huge fan base; like, people fan over this stuff big time. [laughs]LUIZ: Yeah, it was interesting coming from...because when I joined, my first contact with HashiCorp was initially Vagrant. I think everyone goes through the Vagrant status and then Terraform. And I only learned about Nomad when I interviewed. So I didn't know about it before I was...I used to work at IBM, and my team was using Rancher at the time, Rancher 1.x, so like a Kubernetes migration. Afterward, when Rancher 2.0 came out, everything was Kubernetes. So we were like, oh, we might as well use Kubernetes. And then IBM bought Red Hat, which meant that everything became OpenShift. So we had to migrate for a third time. So I kind of dabble in all the different tools out there. And at the time, we were using Terraform for basically everything in our infrastructures. And it kind of became the most fun I had in the day was just playing with Terraform to the point of, okay, maybe I should be working for them since that's where I have the most fun. But during the interview process, it was kind of a generic position, so it was like a systems engineer generic, not a specific team or product. And then, during the interview process, the Nomad team liked my background and picked me. I was like, okay, let's learn about what Nomad is. The first time I just got started, like, Nomad agent-dev, and then you have an environment up and running, ready to use. I'm like, okay, that's very different than what I'm used to.ADRIANA: Yeah, I know. That was kind of my, oh my God, it's the same binary that does everything? What?TED: Luiz, would you mind just since the audience may also not be super familiar with Nomad, just briefly describe what it does and a basic architectural overview.LUIZ: Yeah, sure. That makes sense. So Nomad is a workload orchestrator. So what that means is that it will grab any sort of tasks that you may have, that you want to do, and any type of infrastructure that you have, and it's going to distribute and schedule those tasks into the cloud. So it will kind of, in some sense, abstract away your cluster. You have thousands of machines running. You don't actually care where things run. You just give it a specification, and Nomad will figure out the best place to run and to keep running. So if a machine dies, Nomad will reschedule things and make sure that your specification is always real and it's always as you defined. That's the role of the orchestrator. I guess where Nomad is different is that Nomad is very focused on that task. If you come from a Kubernetes background, you may be aware that Kubernetes does that, plus a lot of other stuff. So it's more feature-complete but also more complex in that sense. There's more to learn about. There's more to understand. And there's a lot more going on when you're just like, I just need a container running; I don't care where or how. So Nomad is very focused on that one task, one job of scheduling things on to other things. Another difference is that since it's focused on the scheduling part, it is generic in terms of the workload. So nowadays, the most common workload are containers, so you can run containers everywhere, but Nomad is not restricted to only that; you can run JAR files; you can run QEMU VMs. You can run Podman containers, whatever. So Nomad has this flexibility in terms of what you want to run and what type of workloads as well, so you can run batch jobs or services that are always running or dispatch jobs. So there are all sorts of different use cases that you can do with Nomad. They are sort of built-in into the core of Nomad. So there's no need for external tools or extra coordination to support things like rolling upgrades, or blue/green deployments, or things like that.ADRIANA: Cool. One of the reasons why we asked you to join us today is that, I guess, a few weeks ago, we were chatting, and you mentioned that you were looking at the possibility of instrumenting the Nomad code with OpenTelemetry, which totally piqued my interest. And as OpenTelemetry lovers, we're like, yay, this is great. So why don't you tell us a little bit more about that aspect that you've been exploring?LUIZ: Sure, yeah. It has been an area of interest for a long time. As I mentioned earlier, I used to be this developer that does operations as well. And at the time, I didn't have a good sense of what it means to instrument an application, what it means to monitor things. My manager did a very good job of getting us large screen TVs with great dashboards and all of that. But they were always, like, doesn't matter how many metrics we have, there's always a problem happening. And every time there was a problem, we didn't know what to do exactly. [chuckles] It was always like, we had the information that we had, but that was never enough to actually solve problems. So now, looking back and then learning about OpenTelemetry and this idea of observability and what it means to have an observable system, it all resonates with me very well because that's the stuff that I wish I had before, and then that's the way things should have been done in the past to understand when an outage happens causing some problems, and things like that. And so now, switching back from this developer operator perspective to this tool builder perspective, I was looking into ways to make past Luiz’s life easier, like, what the tools that I'm developing today could have done to make my life better in the past. And I think observability is one of the major things that we can improve. Because, as I mentioned the description, there is this promise of, like, oh, I don't care how my container runs; I just want it to run. And that works 99% of the time, but that 1% when it doesn't work, that too is completely opaque to you. You can have logs and metrics, but that's not enough to really understand what's going on and what's wrong. So that's where I started becoming interested in this space is just like; how can I make Nomad more transparent and more understandable to people that are using it? And to kind of give that internal view of like, okay, this action triggered these internal operations that generated these internal objects that eventually becomes your container. And so this piece of observability, OpenTelemetry, all of that fits very well into the narrative, both in terms of when we talk about Nomad users, we usually talk about two different personas. So we have the developer persona, which is the group of people that are writing code. They're generating, let's say, a Docker image, and they want to run the Docker image somewhere. And then there's the operator persona, which are the people that are managing the infrastructure, starting the VMs, installing stuff. They're more like managing the infrastructure part of things. When looking into making Nomad more observable, I started looking at these two different personas. What can you offer for all of them? So what the developer cares about, what an operator cares about. And so I started looking into what can I provide to each of them? And then it all comes back to telemetry and what kind of data is more relevant. And so, yeah, from the user perspective, I want to allow them to understand what's going on better, like, what's happening when they run some command or when they do some operation what's going on.But also a little bit more selfish, I also wanted to make my life easier. So when people file bugs, or there's a support ticket, you're always in this situation where one side of the issue is able to collect data but don't necessarily know what data to collect. And then there's the other side of the issue which is us, which is we know what data we need, but we don't have the means to collect it. So there's always this back and forth between look at this metric, what does it say? Give me this log; what kind of information is there? So my hope is that having this common language of telemetry and traces and spans and all of that will give us a more unified conversation and just make our lives easier in terms of supporting our users and also reducing their load as well. Like, if people understand what's going on, they may not need as much to explain what's happening in solving issues. There's the user side of things, and there's also sort of the selfish making our lives easier as well perspective.ADRIANA: I love that, yeah, because there's nothing more frustrating like you get a user ticket, and they're like, blah, blah, blah, it's not working. You're like, oh my God, I don't even know where to start. It's like this terrifying moment.TED: It's really funny. That's actually what you're describing, Luiz, is specifically what got me into distributed tracing, and OpenTracing, and OpenTelemetry, and all of that was having other people operate my software, which in this case happened to be literally the same kind of software, a scheduling system that they're running workloads on. And then they're saying there's a problem. And we need data from them, and the data that we could get without distributed tracing was just kind of a nightmare to dig through because it's like all or nothing. I can't be like, just give me these logs, there's something. It's just like, well, give me a dump of everything off of all 200 machines that you're running, and I'll pore through it over here. And that was just really, really tedious.LUIZ: Yeah, there are a lot of ugly, bad scripts, lots of jq happening to try to correlate all the different logs and all different metrics from different machines that are in different time zones. So it's just sort of this mess. And living in this space of building tools where you're not actually having a SaaS product or you have control over the environment that is running is really challenging because we know nothing about where they're running. We know nothing about their environment. We know nothing about how they're running. So we need to keep probing and asking for more information. And then every question generates five other questions just because in this space where you're shipping tools; you may not necessarily have all the information that we need. We actually purposely don't collect any data from our users. So yeah, that's definitely challenging. And getting the right information that we need to help takes a lot of back and forth.ADRIANA: Yeah, and it's challenging, too, because sometimes you get the bug report of, like, this isn't working. But tell me more, what specifically is not working? You have to keep digging and digging and digging. And the more information you have, the better if you have that information in the form of telemetry data, even more awesome because it makes your life easier.But I also like the point that you made about instrumenting Nomad to cater to both the developer who's deploying their containers to Nomad and having to manage their workloads and the person who's actually operating a Nomad cluster. Because I've been in a position where my former team something would go down, [laughs] and it's like the mad dash. And time is ticking. Because if something's gone down and it's part of critical infrastructure, [laughs] you really want to get to the bottom of that as quickly as possible and hopefully avoid executives breathing down your neck. And so if you have that kind of information, it just makes life so much easier in general.LUIZ: Yeah, when you need the data the most is usually the most critical time because you're like, your system is down. You're having customers asking like, "What's going on?" So that's the moment that you need to be more precise but also the hardest moment to find data because usually, the system is not behaving as expected. So with only logs and metrics to guide you, you rely on your past self to have made good decisions about what to log and when which doesn't always work well. And sometimes, we have to ship a new custom binary for a customer with only one line of code to log something more specific. And then it's like, deploy this binary, and let's see what happens. So these situations where you have no idea what's going on and we need to get more log lines is not a great place to be. Because it's extremely time-consuming to deploy new binary, wait for the problem to happen again, hope that that log line will give us some new information that will answer the question. So yeah, this traditional way of doing monitoring is pretty limit…

    Full show notes at the publisher

    The Evolution of Operating Software with Jason Harley of Honeycomb Sep 27, 2022
    Show notes

    About our guest:Based out of Toronto, Canada, Jason has been working in a variety of roles for the last 20 years in fields ranging from marketing technology to high-frequency finance. He loves helping teams make their systems and platforms more humane to deploy, operate, and reason about. He's currently working at Honeycomb as a Software Engineer building APIs and integrations but started at Honeycomb as a Customer Architect helping customers navigate the sociotechnical aspects of adopting observability practices. Jason is also a HashiCorp Ambassador.Find our guest on:Jason’s LinkedInJason’s TwitterJason’s GitHubFind us on:On Call Me Maybe Podcast TwitterOn Call Me Maybe Podcast LinkedIn PageAdriana’s TwitterAdriana’s LinkedInAdriana’s InstagramAna’s TwitterAna’s LinkedInAna's InstagramShow Links:Honeycomb.ioRedHat Linux 5.2Alias Wavefrontkubectl10x EngineerHashiCorp CommunityTranscript: ANA: Hey y'all. Welcome to On-Call Me Maybe, the podcast about DevOps, SRE, observability principles, on-call, and just about everything in between. Today we're talking to Jason Harley. We're so happy to have you here joining us. Welcome. What is your beverage of choice today, this morning?JASON: It is 10:17 in the morning here in Toronto. So I am drinking boring peppermint tea, two coffees behind me, and some sort of caffeine management regimen. I don't know.ANA: [laughs] There's nothing boring about peppermint tea. It just sounds very soothing and warm.ADRIANA: It's a much more exciting drink than my water. So...ANA: I will say that it's almost a blend of both of y'all. I'm on a Yerba Mate mint tea. So I got a little bit of mint and a lot of caffeine.ADRIANA: Would you believe I'm not a coffee drinker?ANA: No, I cannot believe that.ADRIANA: I do a great disservice to my culture [laughs] because I don't drink coffee, and I don't watch soccer. [laughter] I think I make a pretty bad Brazilian.ANA: Jason, you have many years of experience in the tech industry. And you've been very successful across various companies. So now that you're coming on your first podcast, what an honor. We would love to talk to you about some of the things that you have done. We have listeners wondering how did you get into tech?JASON: That's a great question. I installed Red Hat Linux 5.2 on my family computer when I was in grade 9 and erased all of my parents' stuff in the process.ADRIANA: Oh no. [laughs]JASON: I had no idea what I was doing whatsoever. But I was convinced that this was something possibly associated with some weird hacker identity that one sort of had if you were growing up in rural Canada and predisposed to computers. But I was always fascinated by technology. I was convinced I should do computer science. I only applied to one school and one program. And tech has been really...I wouldn't say I'm a computer scientist by any stretch of the imagination, but tech has been quite kind. I mean, our industry is rife with toxic bullshit. But at the same time, I think there's so much opportunity to grow and learn, to be curious, and connect with people. And all of these skills are immensely transferable. I've done a lot of different roles, and they're all sort of grounded in the same technological experience or principles. But yeah, I moved to Toronto after university and started working at a 3D graphics company, which was called Alias Wavefront or just Alias. They made the 3D graphics software that made Toy Story and Lord of the Rings and that kind of stuff possible.I worked in the business department on infrastructure like email, and routing, and stuff like that for this global company. It was a lot of fun. We got acquired by a very large company, Autodesk. And I promptly left and went to a 30-person startup, which began the next crazy five years of my professional life. I was on call, and all of it...we were transacting money over the internet, and I don't mean like Venmo or Interac; I mean spot foreign exchange. So we made a market, which is to say that we set...didn't ask prices for currency pairs, and we exchanged them. When I left, we were doing north of 10 billion a day in transactional volume. And we did that with a pretty small team of operations folks. My team ran the infrastructure, and we had a team that was specialized in running the company's software. So basically, if you plugged it in, we were in charge of it if it was written in-house, this operations group. But that was my baptism by fire in tech. And I've been in love with startups ever since. ADRIANA: That's so cool.ANA: I think it's awesome to call it baptism in tech. Just like, yep, that was my parents' computer. That's my baptism.JASON: Thankfully, my parents weren't super tech savvy. So they didn't have a lot of stuff on it or anything. But it was definitely the family computer, not my computer. And yeah. [laughs]ADRIANA: That's definitely a good introduction to Linux. I think the first time --JASON: Yeah, I learned that mounts are not pointers is really what happened there. Mounts are not pointers. [laughter] They're the real...block device means something. Who knew?ADRIANA: So, having been in tech for a while, what's been the biggest change that you've seen through your career? I mean, I can speak for myself, like, the tech that I started in when I graduated university is definitely not the same tech [chuckles] that we're in right now. So, yeah, what's your perspective?JASON: I think the biggest change I've seen is automation and by way of programmatic infrastructure. I choose that very deliberately over something like cloud because a lot of enterprises are still able to automate in a big way with things like VMware, for example, or things with Microsoft Hyper-V. But being able to automate and codify configuration and infrastructure, I think, has dramatically changed the tech industry in the last ten years. And I graduated from university in 2004. If I was to tell my new graduate self something to do differently, I think in like 2005; I would have been like, you should really start to learn CFEngine (And there were early remnants of Puppet around that time, I believe.) because I came to that a bit later than I would have liked. I think it was 2012/2013 when I really started picking up Chef at the time, but I think that has been a game changer in terms of converging. We can bring software development practices to infrastructure. I'm working as a software engineer now, like, an actually one. I'm working on APIs and such. But I think from a reliability and a learning and a testing standpoint, that's really changed the game. Developers can come in and participate. We can say what we want about DevOps and the Wall of Confusion. And it's now a proper noun, apparently. I saw a job posting the other day saying that 90% of DevOps like this technology, and that's a wild one to even take apart its actual intention. But that kind of automation I've spent most of the last 8-10 years, last eight years, using AWS in a pretty big way. And the stuff that you can pull together with that kind of tech is truly [chuckles] game-changing compared to the start of my career, to your point.ADRIANA: Yeah, that stuff is trippy. I mean, considering that you can practically at the snap of a finger you've got a VM; you've got a Kubernetes cluster. It's like, what? Kubernetes wasn't even a thing when we were starting out. Docker wasn't a thing. I graduated school in 2001. Java was still nascent. It was the cool kid on the block at the time. So yeah, it's mind-blowing. And now we've got new languages like Go, and then old languages experiencing a resurgence like Python is popular again. I think it's so cool to see that kind of thing.JASON: It's wild, yeah. I graduated university knowing Perl and Java, and I haven't used either of those in a number of years. I've spent most of my days in Golang these days, Golang search engine versus real life, I guess. But yeah, it's been a wild shift. The other part of that that I think is fascinating, and this is to double click on your point about with one call, you can have a cluster of seemingly infinite resources, or you can launch 100 virtual machines. The complexity that has come from that capability is also pretty wild. And this is something we talk a lot about at work. But thinking about operating software is very, very different now when you don't have the database and the app server. And I worked at places where we had the database or at least the finance database, and the customer database or whatever, and then some big, gnarly monolith, whether it was Java and WebSphere or some PHP thing. Those things still exist. They're still totally valid paradigms but managing that complexity has been the wildest part. And that's something that I think our two companies, Lightstep and Honeycomb, do a really good job of talking to people about. I was on call in a very stressful way for five years. I've been on call multiple times since, but that was definitely one of the most foolish things that I've done. It was a lot of fun. I learned a lot. But being able to reason about systems in their complexity, mental models don't work anymore because of these things. I think I would also go back and tell my younger self, and it's something I try and explain to other people, is that mental models are fundamentally broken because the business and the industry demands such high rates of change now, and we achieve those through mind-boggling complexity, but we just talk about them as really abstract concepts. Like, yeah, you just run kubectl or kube C-T-L (Let's not start that debate.), and you push in your new manifest, and then the stuff happens; that's great. But when it breaks down, you need to understand what's going on. In a past role at Honeycomb, I was actually in customer success. I've worked with a lot of our bigger customers through some of both the technical and social challenges of operating complex software. And the hardest part there is to get people to understand that it's okay to not know everything. And in addition to that, you can't know everything. And you just need to accept that and ask questions and be curious. And that is what makes "10x" quote, unquote, engineers; your ability to roll up your sleeves and collaborate and ask questions. Those are such exciting things.ANA: It's very true; those mental models, there's no way to hold them in your head. They're constantly changing in a way that you think you have your architecture diagram mental model built out, but then you have two others go push out a whole new deployment. And that deployment had so many other dependencies, and now it's all changed, and you don't even know about it.JASON: And that's the dangerous bit, right? Operating with a false model is inarguably worse than operating with no model at all. And then there's a lot of ego tied in with that; I mean, we could all do with less. Everybody's got some ego to manage, I truly believe.ANA: It reminds me of like the on-call hero where they just really want to save the day. And you just really want to put your name out there.JASON: I am a recovering on-call hero for sure. But in my past self, we didn't have tools to work through that in the same sort of way. And the only way without observability and much more active documentation and actually talking about empathy across company boundaries, the reason that is such a thing is that we created it. It was a gap that needed to be filled. And without better solutions, I think that's just what happened. And that's not to say it was good; it was terrible. It held companies back. But having lived through that role, it took too much time, too much effort, too much stress. And I think all of those things ended up being functions of our inability to actually collaborate and share those models in a meaningful way. I think we're doing a lot better with that. We still got a long way to go as an industry. But that, to me, I think, is the next big shift. We can make 1,000 computers with a single API call. I think we're getting to the point where we can now start to ask questions and operate more complex infrastructure. We say at Honeycomb that software is a team sport, whether you're SaaS or you're shipping binaries. The medium to be a team is what's really exciting to me these days.ADRIANA: Going back to the on-call hero thing, as you said, it was almost a necessity because that person had the domain knowledge. But with the tooling and the practices that we have now these days, it becomes less and less necessary. So it not only relieves the pressure of the on-call hero, but it also gives them room to grow in different ways than they would have before. Because they're just stuck in this world of domain knowledge, and that's it, that's your life. And it also didn't give room for the more junior people to get up to speed because it's like, get out of the way. Let the big people work on the problem. And now it's like observability has made this an equal opportunity playing field again, back to the whole thing of the team sport.JASON: Exactly. And it was just a vicious, self-reinforcing cycle. It has so many negative drawbacks. But, I mean, a lot has been said about hero culture. I cannot do it justice. But that very short-sighted view with the big dopamine hits and feeling like you saved the day, that's great, but it was not great even in the medium term. And scaling engineering companies and scaling delivery as software teams that is 100% an anti-pattern these days.ADRIANA: Going back to your on-call days, so I guess two questions: how was it before when you were in the early parts of your career where you were doing on-call a lot and things were less evolved practice-wise, tool-wise? And have you been on call more recently now under our newer, more evolved way of...or at least as we all try to practice a more evolved way of doing on call?JASON: Yeah, I've been on call more recently. I have just switched roles in Honeycomb into the engineering organization. And it's a new team. We don't yet have on-call responsibilities, so that's going to change. Engineers at Honeycomb are on call for their area of ownership. And I can honestly say for the first time in a very long time, I'm very comfortable [laughs] accepting that because it's such a different culture, steeped in both good engineering principles and good observability principles. So I'm actually a little bit excited, dare I say, about on-call because it truly is a learning opportunity for how this stuff comes together and how our customers actually use our product. It is not “I'm going to be woken up at 3:00 a.m. at least three times in a seven-day period” to deal with something that is probably trivial. Or, based on past on-call experience, it was either trivial, or the company was literally melting. And I honestly have this memory, being the fickle thing that it is, my memory is that it was like this disk is going to be full, and honestly, a shell script should have been run that compressed a bunch of files or removed a bunch of files. And the other part of it is that all of the production database arrays have just gone offline [laughs] because they're 1,100 miles from you. And you now have to wake up who knows how many people. I mean, that is still a part of a lot of people's jobs. And I don't mean to trivialize that but not knowing even what my week was going to be like. And we talk about unknown unknowns in observabi…

    Full show notes at the publisher

    An OpenTelemetry Journey with Gabriel Fonseca and David Alfonzo of Wavelo Sep 20, 2022
    Show notes

    About our guests:Gabriel is an Observability Engineer at Wavelo, based in Belo Horizonte, Brazil. Prior to his time at Wavelo, he spent several years as a DevOps Engineer working for large Brazilian e-commerce providers.David Alfonzo is the manager of the Platform Solutions team at Wavelo. He is laser-focused on how SRE and software development can make the internet better. During his work journey, he has worked in a variety of roles, from support, web development, sys admin, infrastructure, security, DevOps, and management. He is based in Toronto and enjoys long walks with his wife Durley and the hot summer weather while it lasts!Find our guests on:Gabriel’s LinkedInDavid’s LinkedInFind us on:On Call Me Maybe Podcast TwitterOn Call Me Maybe Podcast LinkedIn PageAdriana’s TwitterAdriana’s LinkedInAdriana’s InstagramAlex’s TwitterAlex’s LinkedInShow Links:OpenTelemetry.ioOpenTelemetry CollectorCloud-Native Observability with OpenTelemetryTucows.comWavelo.comHashiCorpHashiCorp NomadMetricsPrometheusOpenTelemetry Prometheus ReceiverOpenTelmetry StatsD ReceiverOpenTelemetry Jaeger ReceiverOpenTelemetry Zipkin ReceiverSLO (Service Level Objective)OpenTelemetry Protocol (OTLP)CNCFOpenMetricsKubeConAdditional Links:O11ycast Podcast Episode: Ep. #54, Cloud Native Observability with Alex Boten of LightstepAlex Boten on MediumDavid Alfonzo on MediumTranscript:ADRIANA: Hey, everyone. Welcome to On-Call Me Maybe. I am your host, Adriana Villela. And today, I have a special guest host with me, Alex Boten. Alex, why don't you talk about yourself a little bit here?ALEX: Hi, everyone. I'm Alex. I'm an OpenTelemetry collector, contributor, and maintainer, and I'm also a contributor to some other projects within OpenTelemetry. I'm a senior staff software engineer at Lightstep. And I'm also the author of Cloud-Native Observability with OpenTelemetry.ADRIANA: Awesome. And today, Alex and I are going to talk to two former colleagues of mine from Tucows. We have Gabriel Fonseca, and we have David Alfonso. So, guys, why don't you introduce yourself? So let's start with Gabriel.GABRIEL: Hey. Hello, everyone. I'm based in Brazil, and I have been working with Tucows in the observability team since last November. And we have been trying to move to OpenTelemetry inside Tucows and get everything it has to offer.DAVID: Nice. I'm David Alfonso, and I’m actually based in Canada and working closely with the observability team as well. I'm the manager for platform solutions within Wavelo Tucows. I'm pretty much as well, just trying to get in the platform, specifically trying to get OpenTelemetry working and some adventures that we have with that.ADRIANA: Cool, awesome. So the reason why we have both David and Gabriel here is that full disclosure, they both used to work for me. I used to be their manager. So David now has my old position as manager of the platform solutions team. And I also managed the observability practices team at Tucows Wavelo, and Gabriel was one of my hires. And as part of our mission to bring observability into the organization, we wanted to go the OTel route. David and Gabriel are going to talk about some of their adventures in basically bringing OpenTelemetry to the organization and specifically around the OTel collector. So why don't we start with the OTel collector? What were your experiences in running the collector? And specifically, what was some of the architecture that you guys had to deal with? Because you guys aren't a Kubernetes shop like most of the world out there; you guys are a Hashi shop. So what were some of the challenges in running the collector?GABRIEL: Well, I think maybe the first one was how to deploy that in Nomad and figure out all the small things that have to change, but Adriana got it perfect for us in the first place. So we are still running these in pre-prod. We have tested this with several teams, the auto collector on our site. David's team is one of them that is sending a lot of metrics to us from all the Nomad nodes we have. Maybe you can talk more about that, David; how does that work for you?DAVID: We pretty much tried it all, to be honest. [laughs] One of the things that we did is like you said, we are a HashiCorp shop. So we use Nomad, Consul, both, you name it. Anything HashiCorp, we have our fingers in it kind of deal. So part of that is how do we get metrics from the right places in the right ways? So the goal for us was going directly from the host all the way to our vendor. But now, in order to do that, we need to do a few steps. We started with...on the host specifically, we collect the host metrics. And we name the host metric right in the beginning. We figured out the hard way, the closer you are to the data, the easier it is to work with collector. You can name things, for example, directly on the collector, and you can set all this stuff directly that way and then send it to another collector if you may.In our case, we have, you can call it a gateway collector or a central collector for Nomad specifically. So we will send from the host to this middle layer collector, to Gabriel's team's collector, and then from there, ship it to the vendor. And we find that that works really well and just for a few things. The first one was it was very simple to set up. The con part of that is you have to set up in multiple places. And we use a configuration management which makes life way easier to set everything up.The other piece why we decided going directly to the host instead of, for example, scrapping directly from a Prometheus telemetry and directly from the host; we decided to go directly to the host because, one, we can name things directly on the host. Like if I'm going to say Nomad client one, then I know that it is coming directly from the collector Nomad client one. And so I don't have to do shenanigans or complex setups and stuff like that later on in the other collectors. So that's some of the things that we did in order to get all the way down to our vendor.ADRIANA: One of the things that I remember from back in the day is that when you started ingesting metrics from the Hashi stack, you guys were using a really early version of the Prometheus receiver. So what were some of the challenges around that? Because my dream when I was there was basically to get rid of Prometheus. So I'm like, I really wanted to make that work. [laughs] But what were some of the challenges, and now that I think that receiver is a little bit more mature, how is today's receiver different from last year's receiver?DAVID: Full disclosure, we were using Prometheus before, like Adriana mentioned. For us, it was like, how do we use the collector to use the exact thing that they were using before, and we wanted to grab that so just to escape Prometheus out of the picture, and then just put the collector right in between kind of deal. And now the problem with that is if you're deploying the collector the way that we were doing it, we were using TerraForm, like we said, by HashiCorp, and then deploying the collector using Prometheus, then you need to use multiple variables on top of each other, and it becomes a mess. It was very challenging to get it working, especially in the early stages of the Prometheus telemetry collector receiver kind of deal. And because in OpenTelemetry stuff as well, we were using...we were going directly for the Prometheus receiver itself. So that's some of the challenges we run into. I don't know, Gabriel, did you run into anything specifically in your site?GABRIEL: I remember at some point, we had some issues with the collector breaking because of some misconfiguration from the receiver. But I'm not exactly sure what that was. But Prometheus receiver has replaced almost all Prometheus, I think, in the company. I'm not aware of any...that's probably Prometheus running somewhere because there's a lot of stuff running there. But I'm not aware of anyone that is using that part. So we got rid of most of that, at least. And also, on the metric side, another thing that we use a lot is StatsD metrics format. And for that, we also set up the StatsD receiver. And luckily, the OTel guys added the label support to it from log StatsD. This was really great because it enabled us to remove a lot of garbage metric that we had. We still have some, but we're working on it. So that named the metrics and replaced the names by labels and so on, so restructuring that instead of having huge names. And it has been working very well so far. Besides the Prometheus receiver, we are also using the StatsD broadly.DAVID: Alex, I'm wondering, how did you guys, like, did you run into something similar to this? And how do you guys fix it?ALEX: Unfortunately, or maybe, fortunately, I don't know, like every other large-scale deployment internally at Lightstep, we also run multiple metrics solutions. And I think we're still in the process of migrating some of them. But we're using the StatsD collector. We're using the StatsD receiver, the Prometheus receiver. And we're; also, I think we're just in the end of removing another component that we called...it's basically a metric proxy for StatsD. It's taking some time. I think the collector has matured significantly in the past year. And so, there was a lot of hesitation at the very beginning to migrate on to the collector because it was still very much in development. But now that it has support for so many different formats, it's become a lot easier to manage.ADRIANA: So, Alex, is there a specific component of the collector that you work on?ALEX: No, just a little bit everywhere. I'm a maintainer on the core collector, which is the area that I'd like to focus as much as possible because there are only a handful of contributors there. But the contrib repository is the busiest OpenTelemetry repo by a magnitude of I think it's 5X over the next one. So there are a lot of people that are really interested in getting their PRs, and so I’m trying to help as much as possible there.ADRIANA: Cool. That's awesome. Maybe for folks listening in who aren't familiar with the two repos, can you explain the difference between the core collector repo and the contrib repo?ALEX: Yeah, the core collector repository is where the functionality that's maintained by OpenTelemetry lives. So anything that's open source so, like the Jaeger receivers, the Jaeger components, or Zipkin components, or Prometheus components, will eventually move to the OpenTelemetry core repository. And the contrib repository is a little bit more open as to what it accepts, so it accepts contribution from vendors, contributions from individual users. Basically, if you're looking for something that's outside the core supported formats in OpenTelemetry, you'll find it in the contrib repo.ADRIANA: Awesome. And I'll put in a plug in the middle here saying to anyone who wants to contribute to OpenTelemetry; the OTel folks are always looking for contributions, so don't be shy. You can get started out anywhere, even with the docs. That's where I got started. And the collector is written in Go, right? ALEX: It is, yep.ADRIANA: Yeah. So if you're a Go pro, then consider contributing to the collector and collector contrib. Cool. I guess another thing that I wanted to touch on with regards to the collector is I know when I left the observability team at Tucows, you guys were starting to get the collector ready to run as a gateway, as a centralized gateway for basically ingesting OTel data from various applications. So maybe if, Gabriel, you could describe some of the things that you needed to do to make, I guess, the OTel collector gateway more productionalized, some of the considerations that you had to make to make that happen, and where you guys are at with that now.GABRIEL: Sure. One of the things we had to do was moving some stuff more close to the teams, like the Prometheus receiver, for example, because most of the collector is stateless. But some components are not like the Prometheus receiver. And what it does is that if you have more than one copy of the collector running with the same configuration, you have your host scrape it more than once, and then you have to duplicate that later. So one thing we decided was not to have those Prometheus rules on the gateway site, so things just pass through there. We add some tags. We can do filtering and things like that. But we do not have any stateful module there. Other than that, I think everything else was almost straightforward. We had to take care of some memory configurations. There are some balance configurations for the collector, memory limiters, batch processors, things like that to make sure it doesn't break. Also, we are working on trying to scale it using the metrics that the collector generates. So we generate some metrics about the...so the number of metrics coming in, going out, and so on by pipelines and whatnot. So we are working on scaling that based on these metrics. Yeah, we have also set some SLOs based on those metrics. We have not moved everything yet to the gateway, to our collector gateway, so most people are still using other ways of sending things to the backend. But yeah, we are on the move also working on things not related to the collector, so enabling the developers to be able to instrument properly their code and sending things there. Just a handful of people are using OTel natively. So we are working on providing some shared libraries for people to do that, and so on. So it has been a lot of political work, not just tech work on that [laughs] in convincing people to adopt that and move to our gateway and so on. But I think slowly; we are able to do that. So more people are interested in that and making good questions about observability. And we can see that the observability is really improving in there. So I think we are happy with this. ADRIANA: That's awesome. And what kinds of security considerations do you have to make for being able to run the collector in gateway mode? Do you need to install any certificates? Do you have any SSL-type of considerations when running the collector when you get it ready basically for primetime?GABRIEL: Well, we are still running this all in pre-rod. So we did not set up any certs yet. But the plan is to have this in place when we go to prod. One thing that we wanted to get to work in, but we were not able yet, is how to enable teams to send their own API keys and get these passed through the collector and sent to the backend. So, for example, if we want to send to a different storage in the backend, some folks allow you to send a header and the request, and they will do that for you. But now, we are not able to get this from the teams. We can only set these in the gateway itself. So this is something that we wanted to get ready for prod. Maybe Alex has some insight. [laughs]ALEX: I was going to ask you; I think there's an open issue around this, if I'm not mistaken. I seem to remember seeing being able to forward on a header as an issue somewhere.GABRIEL: Yeah, I think so. Also, making questions for you, Alex, I've heard that the guys in OTel are working on a query language for creating metrics from spans and things like that, right?ALEX: Right, the telemetry query language that's being currently developed in the transform processor. GABRIEL: Yeah, that's pretty cool also. ALEX: Yeah, ideally, at some point, the transform processor will allow the collector to sunset some of the other processors that have been created. I think one of the main things that I came into when I first looked at the collector was there are so many different processors. It's hard to know which o…

    Full show notes at the publisher

    How to Rock at SRE with Liz Fong-Jones of Honeycomb Sep 13, 2022
    Show notes

    About our guest:Liz Fong-Jones is a developer advocate, labor and ethics organizer, and Site Reliability Engineer (SRE) with 17+ years of experience. She is the Principal Developer Advocate at Honeycomb for the SRE and Observability communities and previously was an SRE working on products ranging from the Google Cloud Load Balancer to Google Flights.She lives in Vancouver, BC, with her wife Elly, partners, and a Samoyed/Golden Retriever mix, and in Sydney, NSW. She plays classical piano, leads an EVE Online alliance, and advocates for transgender rights.Find our guest on:Liz’s TwitterLiz’s LinkedInLiz’s websiteFind us on:On Call Me Maybe Podcast TwitterOn Call Me Maybe Podcast LinkedIn PageAdriana’s TwitterAdriana’s LinkedInAdriana’s InstagramAna’s TwitterAna’s LinkedInAna's InstagramShow Links:Honeycomb.ioObservability Engineering (free e-book download - limited time only)Observability Engineering (dead tree edition)Cloud-Native Observability with OpenTelemetryObservability Engineering book signing party - San FranciscoContinuous ProfilingOPEX (operating expense)Jeli.ioNora Jones (Jeli)Unknown Unknowns DR (Disaster Recovery)Multi-CloudGoogle SRE Book (aka “SRE Bible”)Emily Freeman Justin Garrison Twitter Space - DevOps vs SRE with Emily and JustinSLOs (Service Level Objectives)Implementing Service Level Objectives - Alex HidalgoGoogle Site Reliability WorkbookGoogle Seeking SRE BookAmy Tobey’s SRE modelSREconAdditional Links:Trans LifelineAdriana on O11ycastAna Margarita on O11ycastTranscript:ADRIANA: Hey, y'all, welcome to On-Call Me Maybe, the podcast about DevOps, SRE, observability principles, on-call, and everything in between. Today we are talking to Liz Fong-Jones, who is the Principal Developer Advocate at Honeycomb, and she is also on the OTEL Governance Committee. Welcome.LIZ: Thank you for having me on.ADRIANA: So first things first, we like to ask all of our guests, what are you drinking?LIZ: I am drinking a homemade mocha, just a French-pressed coffee and some hot chocolate powder dunked in it. I used to live closer to a coffee shop where I could get real espresso, but now I live a little bit more into the boonies. So now I have to add powdered hot chocolate to my coffee to make a mocha.ADRIANA: [laughs] Awesome. Hey, whatever works. How about you, Ana, what do you have?ANA: I'm just doing classic cold brew with oat milk. And I added vanilla flavor to it to give myself a little different taste this Monday. But that's about it. Pretty simple. What about you, Adriana?ADRIANA: I have a bubble tea today. I'm a huge bubble tea junkie. [laughs]LIZ: Oh, right, because it's noon over there. Unfortunately, bubble tea places...I would love to have bubble tea in the morning. If someone opened bubble tea that was open at 8:00 a.m., they would have a monopoly on the market.ADRIANA: Yes. LIZ: Because I don't want to wait until noon to have my bubble tea. But all the bubble tea shops open at noon.ADRIANA: It's so true, and it's so sad. Sometimes I actually will pre-buy my bubble tea as long as it doesn't have tapioca because I think after about an hour, the tapioca goes really nasty. But I'll get it...if you get it with coconut jelly or basil seeds, it'll survive the night, so then I can have it handy for the next day. [laughs]LIZ: Business ideas brought to you by On-Call Me Maybe.ADRIANA: That's right. That's right. So, Liz, you just came out with a book. So, why don't you tell us a little bit about that?LIZ: Yeah. So my colleagues and I, Charity Majors, George Miranda, and myself, have just published Observability Engineering. It came out in May, and print copies have been available since June. And the book is about the why of observability. We do go into some of the how but I think we're kind of orienting people around how is observability different from monitoring. When should you use observability? How do you introduce it into your organization? What workflows does it enable? There is one chapter in it about OpenTelemetry. But we actually had your colleague, Alex Boten, at Lightstep also publish an entire book-length volume on Observability with OpenTelemetry. So the two books kind of nicely pair up with each other because Alex's book is more about the how and our book is more about the why. And actually, it turns out that people buy them together. Who knew? If you go and look at the Amazon listings, people just buy them as a bundle set, and I think that's really cool. ADRIANA: That's awesome.ANA: And if folks want to take a read for free, as I think I saw on Twitter, there's a way.LIZ: That's exactly correct. Limited time only, you can go to a page on info.honeycomb.io (We'll put it in the show notes.) to get a free copy of the book in PDF format. And there is no paywall or no registration wall or anything. We don't ask for an email address. But if you happen to want a dead tree copy of it after reading the PDF version, certainly, our editors at O'Reilly would appreciate you tossing them some money. ADRIANA: And I remember talking to Charity recently, and she said that if you order a print copy and let her know about it, she'll send you some stickers to bling up your book cover.LIZ: Yeah, that's exactly correct. So I'm handling distribution in Canada and Australia, and she's handling U.S. distribution. But yeah, you can put a sparkle unicorn tail on the maned wolf that's on the cover. I asked for booties, like, Ruby slipper booties for the wolf [laughter], in case you want to stay a little bit historically accurate to what a maned Wolf is and hypothesize that you could put booties on it.ADRIANA: I love it.ANA: That is amazing.ADRIANA: It's almost like a personalized book signing. Send stickers, and that shows that you've got, like, it's the signature from the authors almost, which I think is such a cool idea.LIZ: Yeah, we also do book signings, though. There is a book release party in San Francisco in the middle of September if you're listening to this before the middle of September. So if you stop by the book signing party, we'd be happy to give you stickers in person as well as to actually autograph your book with a hand-written personalized autograph.ANA: Oh, I actually might need to get myself to that party, is what it sounds like. [laughs]ADRIANA: Yeah, it's close to home for you. [laughs]ANA: Yeah, I'm still in the Bay Area, so definitely might be doing that.ADRIANA: That's awesome. It's interesting, with regard to the book on Observability Engineering; I think it's still so important to keep up the conversation on what is observability. Because I don't think enough people understand it fully. So I think it's good to keep hammering the point home.LIZ: Yeah, exactly. One of the things that we frequently wind up saying and wind up I think everyone on this podcast is in agreement with, right? Like, it's about the outcomes that you achieve. It's not about the various signal types. It's about how do we actually debug unknown behavior in our systems? And it doesn't really matter what combination signals you use to achieve that. For instance, one thing that I like to emphasize is that I spent almost over a decade at Google. And in my over a decade at Google, we didn't actually, at least for the first seven, eight years of that time at Google, we didn't really use traces very much. But we had very sophisticated metrics slicing systems that would enable you to join multiple different metrics together, would enable you to do some of the higher cardinality things with metrics that you might not be able to do in a more primitive system. So I would argue that we had observability at Google without collecting all of the various signal types.ANA: That's really cool work. And I also wanted to mention, like, congrats on the book. Now that that's finished, what are you most excited about to be working on right now?LIZ: So it's definitely been really nice to get back to practitioner stuff and prototyping rather than writing. I tend to oscillate between prototyping and writing about it. But the book was kind of this three-year-long project, not necessarily all three years spent writing continuously. But towards the end, it was a little bit of a slog. I wound up not getting to do very much hands-on engineering. So what I'm up to right now is really, really looking at what can continuous profiling do for us. How can it benefit us? And I'm not going to call it a fourth pillar because I don't believe in pillars of observability. [laughter] But I do think that there are some interesting applications of continuous profiling that are not the ones that the people who came up with had originally envisioned. Specifically, people talk about continuous profiling in the context of, you know, oh, we're here to save you 5% or 10% or whatever off of your data center bill. To me, that's missing the point. Yes, OPEX is a concern, but I care about optimizing user visible response times. That's why we do tracing. At a certain point, I know that tracing breaks down. You're not going to put a trace span around every single function call. But in our profiling, where you can actually capture every function call that takes longer than ten milliseconds, 100 milliseconds, you can start to fill in some of the gaps that you wouldn't be able to seal with tracing.ANA: That's actually really cool. I have zero background in the profiling stuff, but definitely would make a lot of sense when we're thinking about the scale that we have our applications, and the users that we have tuning in, and the stuff that's still missing.LIZ: It's so powerful, and yet it's something that, at the moment, is for power users only. And I think that there are ways that every software developer can benefit from it; they just don't know it yet. So that's what I'm spending my time exploring. One of the blog posts I'm working on that by the time you hear this might actually be published is a blog post on how we sped up ingestion into Honeycomb, sped up the ingestion processes in Honeycomb by 50%, which then gave us 10% more capacity to run queries in Honeycomb, which does not directly translate to 10% less latency, but it does translate to in general Honeycomb users queries running faster. So that's something that I found with profiling that I would not have been able to find any other way.ANA: It's actually really cool to hear about some of that work because I know that part of the IC work that you get to do is also help improve the SRE practices of Honeycomb. I know that you were recently on stage at AWS talking about some of the improvements you were doing with some of the migrations that you were doing; I think it was Graviton.LIZ: Yeah, exactly. So part of what's fun about being a principal-level engineer for me is getting to go on wild goose chases. And sometimes, you come back with a domesticated tamed goose. I think that one of those wild goose chases was started two years ago, two and a half years ago, when I got us to start trying out the AWS Graviton Arm architecture-based processors. And that led to essentially a 50% reduction in Honeycomb's OPEX or at least in our compute OPEX. Obviously, paying people salaries, that's something that doesn't get affected by your processor architecture. ANA: [chuckles]LIZ: So getting to explore some of those things without taking cycles away from the product engineering teams and then when I find something worthwhile bringing it back into the org and figuring out how do we leverage this as best as possible? And funnily enough, to wrap this back around to OTel, OTel was actually one of those wild goose explorations. It was a, you know, hey, the Honeycomb Beelines are working more or less well enough, but there's this OTel thing that's new that's combining OpenCensus and OpenTracing. Should we maybe develop an exporter?ANA: Super congrats on being able to chase some wild gooses because we know that in general, the job in DevRel and the work that we do is a lot of context switching, so to be able to get enough hours in a cycle to be like, oh no, I'm actually going to be dedicated to this topic and actually do some deep dives.LIZ: Yep, do the deep dives and then write about them. That's kind of the other piece about it is if you're not communicating about what you're doing, is it ever...or it's a little bit challenging to show the impact that you're having.ADRIANA: It's so true. That's one of the things I actually like about DevRel is being able to prototype because that's like one of my favorite things is take a problem that looks interesting, try to solve it, and then write about it. And, I don't know, I find it so, so satisfying. ANA: I agree. I definitely wish I did more of it currently. So I'll hopefully find more time to carve into that.LIZ: Right. And given that we work with the SRE and DevOps communities, one of the most impactful ways that I found is to sit in on on-call rotations. You don't necessarily have to even be on call, on call, but at least shadow people and see what's going on and see what patterns you can generalize out, what incident response practices that you think the world should know about.ANA: You actually got me right to one of the questions that I wanted to ask; what is on-call at Honeycomb like?LIZ: Yeah, it's something that is constantly evolving because we have grown from having...when I started at Honeycomb, we had ten engineers, and now we have about 40 engineers. And therefore, we have had to split on-call rotations. So now there are actually three on-call rotations. There is the platform engineering on call, and there's the product engineering on call. And then there are the integrations on duty. So those are kind of our division responsibilities. So anything to do with client SDKs, anything that runs on the client's premises that is integrations on call. Anything to do with the product UI is product on call, and anything to do with data ingest and querying goes to the platform on call. So that's loosely how it's divided for now. But we're already thinking about what the next flood of that might look like. So it's definitely interesting to be hiring people and then figuring out how to onboard them onto the existing rotations and also figuring out how much scope is the right amount of scope for one person to keep in their head. Certainly, it used to be the case that we would just have one on-call rotation. And that person had to be responsible for any integration question, and any JavaScript question, and any platform question. And you might try to pair up to make sure that there's usually a front-end-y and back-end-y person on call as primary or secondary at any time, but now that's actually formal. But yeah, the general philosophy is if you ship code at Honeycomb, you are responsible at least part of the time for the consequences of that. And that means that you both watch it as it goes out, even if you're not the person officially on call. And you also take a turn being the person on call for your particular area of responsibility and helping other people shepherd their changes out.ADRIANA: And how has that gone? Do folks at Honeycomb respond well to that? Because I know that that's something that you and Charity are always talking about on Twitter. And I think it's such a great idea. I mean, I think it makes people more responsible for their code rather than, oh, I'm done with it. It's someone else's problem now. What has the reaction been?LIZ: I think people really…

    Full show notes at the publisher

    Previous 1 2 3 4 Next

    Related Podcasts

    Reply All

    1

    Reply All Games & Hobbies
    Inside VR & AR

    2

    Inside VR & AR Gadgets
    Note to Self

    3

    Note to Self News
    BrainStuff

    4

    BrainStuff Natural Sciences
    This Week in Tech (Audio)

    5

    This Week in Tech (Audio) News
    Hands-On Tech (Audio)

    6

    Hands-On Tech (Audio) Technology
    footer-logo

    Contact Us

    Toll Free: 844-670-7747

    Links

    • Home
    • Top Charts
    • Networks
    • Apps
    • Independents Podcasts
    • Podcast Advertising
    • Podcast News
    • Contact Us
    • About Us
    • Analytics & Insights

    Stay Connected

      Privacy, Terms of Use & Our Code of Ethics Protecting Content Creators Copyrights