Webinar Host: We are pleased to welcome you to today's webinar, Generative AI, Lessons from the Frontline. Before we begin, we wanted to cover a few housekeeping items. For the best feeling experience, we recommend logging in on a computer and closing anything running in the background that can cause connection issues. You can send the questions within the Q&A box at any time during the presentation. We'll do our best to respond to your questions, but if we run out of time, we will connect with you after the webinar. At the bottom of your screen are several engagement widgets that you can use. All of these can be moved, resized, or minimized, so feel free to customize your desktop space as you see fit. I will now turn it over to today's speakers.
Roger Burkart: Hello, I'm Roger Burkart. I'm the Chief Technology Officer for Capital Markets and Co-Head of AI at Broadridge Financial Services. Thank you for joining us today for our discussion on generative AI, lessons from the front line. I've been using AI for over 20 years, starting in my role as CTO at New York Stock Exchange, where we used it for market surveillance, both in transaction analysis and in social media, and then in many other CTO and fintech CEO roles since then. I've seen some great big wins. I've also seen an awful lot of AI pilots that didn't really get to material value. Pilots have stayed on the shelf. All that said, I'm super excited about the step function increase in value generation that generative AI adds to the existing set of use cases, such as predictive and prescriptive AI. As all of you on this call know, I'm sure, gen AI is just a year old, really, in terms of the public consciousness. It exploded into the world with chat GPT adoption rising to 100 million users in just a couple of months. And many, many people would say it's just a year old, even though it's been around for a number of years. I've used transformer models back five years from now. I think the big thing is that the chat interface led to a radical democratization of the technology, a bit like the Internet did for access to data. And there's a tremendous interest in applying it in financial services, which I hear from the C-level clients that Broadridge serves. That said, despite that interest, many financial service firms are grappling with how to capture value while ensuring safety, accuracy, and compliance in a very regulated industry. And so there are not that many organizations that have deployed in production client-facing applications. And this is what we'll try and address today. We're going to cover which use cases are adding value. what benefits they're creating, and then what practical steps can you take to successfully implement generative AI in your organization? And finally, most critically, how to organize for AI, how to deal with all the talent questions. So I'm very, very privileged to be joined by two esteemed panelists, Joseph Lowe from Broadridge and Brent Swidler from AWS. They bring a rare real-world experience in delivering value with generative AI solutions. So let's get rolling. Let me turn first to Brent and ask him to introduce himself and his product area that he's responsible for at AWS.
Brent Swidler: Sure. Thanks, Roger. Hey, everyone. My name is Brent Swidler. I'm a principal product manager at AWS. I specifically work on Bedrock. I've been with AWS now for a little over four years. I'm all in the AI space and their language services. And I've been on Bedrock now for a little over 18 months. For those of you who aren't familiar with Bedrock, Amazon Bedrock is a managed offering to access a wide variety of foundation models from a number of different providers, including from Amazon. So you're able to go to Amazon Bedrock and choose from different models. There's the first-party models from Amazon called Amazon Titan, which I've been spending the majority of my time over the past few months working on, as well as from different providers such as Anthropic, AI21, Cohere, stable diffusion for text-to-image models, and then also Meta's Llama family of models as well. You can choose from any one of those models, use them out of the box, or customize them to your needs, and then deploy them into any application that you want to use. So happy to be here. Been in the AI space now for almost seven years, four of which at AWS.
Roger Burkart: And over to Joseph.
Joseph Lowe: All right. Hey, everyone. I'm Joseph Lowe, head of enterprise platforms at Broadridge, responsible for our enterprise-wide capabilities. At Broadridge, we're excited to be developing enterprise-wide AI, you know, capabilities that allow any of our businesses and products to use AI. Today, one of the things that our team has been working on closely with the LTX team is delivering bond GPT, which is what we believe to be the first large language model powered application for fixed income, allowing bond traders, portfolio managers, and anyone else in the ecosystem to be able to use natural language to be able to surface pre-trade insights and workflows using generated AI. So excited to be here today, Roger.
Roger Burkart: Thank you, Joseph. Your practical experience at the front line is going to be very, very valuable. So I'm looking forward to the conversation. What we thought we'd do next is we would just get a sort of sense of the mood of the room, if you like, the virtual room, with a poll. And so I just ask you to respond to the poll that you should see on your screen, which indicates just how often do you use generative AI tools? We're broaders, we're big believers that this is an area where you learn by doing. So I think it'll be interesting to see, you know, who's been getting out there and using generative AI tools, either in your personal, with personal access or indeed work access. So while you're voting on that, I'm going to turn to Joseph again as someone who spends his whole life serving financial services clients and ask you just to spend a couple of minutes telling us about what kind of use cases, admittedly in the early going, but in the early going here, what kind of use cases are we seeing there's interest for in financial services? I can't hear you. Joseph, I think you're maybe on mute. Sorry, Roger. This is my problem.
Joseph Lowe: No, sorry about that. So when it comes to financial services, people are really, really excited about being able to use generative AI, chatbot-based tools to deliver knowledge to their users. So we're seeing a lot of that when it comes to customer service. when it comes to operations users, even when it comes to the frontline, people want to be able to access all the information that firms have already collected, documentation, procedures, and guides like that, have it all surfaced up through a chatbot. This is also really becoming very common, and a lot of people are talking about being able to also now do that when it comes to user manuals and helping financial services firms, their operations teams, their advisors be able to use the software, the capabilities better. Another example of use cases that are really picking up is where there's a desire to aggregate complex data and disparate data sources and be able to kind of bring them together using natural language, making it so that insights from the data can be gained from natural language as opposed to needing to go to a data scientist, needing to go to your data analyst to do that. And the third one I'd probably highlight is there's a lot of desire to personalize the customer experience. We think that generative AI's capability to summarize content, be able to remix content with a certain tone, a certain vibe, and be able to kind of aggregate content from different sources into a vibe, something that people are really interested in.
Brent Swidler: Joseph, one of the things that I hear in the financial services space particularly is that, and maybe you can substantiate this or not, but that folks are less reliant on using the models out of the box for any sort of financial strategy or disciplines, right? Because they're more looking at how they're leveraging their own proprietary ways of you know, making decisions and having the models being able to read and write from their own proprietary sources as opposed to just, you know, blanket out-of-the-box usage.
Joseph Lowe: A thousand percent, Brent. Couldn't agree more. You know, I think in financial services, there's no real interest in using out-of-the-box, hey, whatever ChatGPD was trained on. You know, I think it's more about using the large language models to understand the language that's already been written within the firms. Yeah. Another area that I'd add is firms are really hesitant to use anything out of the box because they tend to not understand the firm's compliance or regulatory kind of environment that they're in. And so a lot of people are looking at how we can apply compliance engines, compliance layers to the output that's happening. Interesting.
Roger Burkart: Super. So we've got on the screen now a little temperature check of the audience. So almost two-thirds of you are using generative AI tools either daily or weekly. Some folks have yet to dip that toe in the water, and I'd encourage you to do it because I think we learn by doing in this space. And then we have our 17% for using it monthly. So, Brent, you deal with a lot of different organizations. Does this surprise you, or is this pretty similar to the normal distribution you see elsewhere?
Brent Swidler: The number of people that said never is surprising to me. I would think that that would be a lot lower, you know, more in the daily, weekly, monthly section than never. It's rare that I come across folks that maybe it's just bias of like the space that I'm in and the conversations that I have with specific folks. But I feel like it's rare that I come across, you know, individuals that have never had access or don't use Gen AI tools on some sort of regular cadence.
Roger Burkart: Well, I guess maybe those folks are coming to the webinar to understand how to go.
Joseph Lowe: Yeah, I would add that perhaps they don't know that they're using Gen AI tools. You know, with iOS 17, you know, that's that your autocorrect is all now Gen AI, which is amazing. And I'm surprised by this as well. But, you know. I think in the tech bubble and the fintech bubble, sometimes it feels like that's all anyone ever will talk about. And so it's actually a reminder that there's a lot more to do when it comes to explaining the benefit of these tools to everyone.
Roger Burkart: Super. So we're going to move forward now to another question for our panel. Some firms are holding back from implementing generative AI tools because of concerns around safety. accuracy, and then all that financial services compliance and regulation I talked about. So I really wanted to ask each of the panelists, starting with Brandt, to talk a little about what kind of practical advice can you give about how to overcome these issues to actually implement generative AI specifically in a responsible way, and along the way, get your risk management colleagues comfortable? Yeah.
Brent Swidler: Yeah, so I'd start off by saying that the models are probabilistic by nature, so there's always some inherent risk in what they generate or the types of content that they can output. The way that we end up seeing a lot of customers go about this analysis is looking at different use cases along a spectrum of applications based on how they want to go about implementation and what sort of level of risk or communications that they want to have with their end customers or however they want to implement the solutions. We see a lot of folks starting out with just building internal functionality that helps them reduce risk because they're basically giving access to their internal employees to test out different use cases and just generally play around to see what they can uncover. Now on the other opposite end of the spectrum is when you have something like a consumer facing or a customer facing like a customer service facing chat bot, which the model is now somehow representing you as a business to your end customers, which is inherently a lot more, you know, has a lot more controls that need to be put in place to make sure that such models end up communicating in the way that you want them to to your end customers. The way that we see a lot of folks overcoming a lot of this is looking at this as a spectrum and starting out with the more simplistic ones that allow them to test out different things with a lot more controls as they move towards things like use cases that are a lot harder to control the end users interactions with the models.
Roger Burkart: Can you share any examples from AWS? Because I guess your users also, where have you chose to start?
Brent Swidler: So yeah, we definitely give access to all employees to just a blanket playground where they can interact with any of the models, both to give them the ability. We see two things come out from that. One is that people are better writers. They're able to expedite a lot of their... They're able to be a lot more productive in certain areas, whether it's like code generation or... you know, getting content or rewriting content or anything like that. But then we also see them test different use cases for whatever product or team that they're on. So for those applications, we end up, you know, by just giving people a playground to, you know, play around in, they both get the productivity improvements that they want from just, you know, being able to leverage the tools, but also uncover use cases that get brought up into what we have as like PR FAQs and eventually into functionality into different products. Some of the interesting things you might see if you shop on Amazon, as an example, I don't work on this directly, but I can see it as a consumer, that on the more easier end of the spectrum, from my perspective, is things like summarizing reviews, which you don't have any user interaction directly with, which makes the risk of something like that a little less than if you had a full functionality where someone could just interact with a model and ask it inappropriate stuff or try to manipulate it in a certain way. So that's on the easier end of the spectrum. You can see that that's something that can be deployed with a lot more guardrails in place because you have full control over what the model outputs and how it's used.
Roger Burkart: Super. I'm turning to Joseph. You actually had experience of bringing products to market for those easy to please people, fixed income traders, and getting it through all the legal and compliance hurdles of doing that. What have you learned along the way that you can share as to how to overcome concerns about safety, accuracy and compliance?
Joseph Lowe: Absolutely. I mean, I think the first thing, and this is actually not AI specific even, is that you have to know your user. Know who you're serving and what their expectations are and what regulatory environment they live in and what regulatory environment we live in. And so starting with that, when we were delivering bond GPT, it was, okay, we had to make sure that our application wasn't going to give advice. It wasn't going to opine on what was a good bond or what's the most appropriate bond for any given situation. As we started uncovering those rocks and understanding what our needs of our users were, we started at the same time learning by doing, understanding that in a non-deterministic model, we had to build guardrails for what people could type, what people could ask of the application, but also the types of responses the application could give. A good example, actually, in our prep, Brad and I were just kind of talking about how some of that stuff works. And a good example is that for Bond GPT, for example, we have a classifier that sits ahead of the whole application that determines whether what the user is asking is a question we want to answer or whether our model is capable of answering. And it goes from there. And so every step of the way, you have to ask yourself, what's the user's expectation? What's the environment they live in? And then start from a place of no trust. When it comes to LLM, you want to make sure that you assume you can't trust what the user is going to say. And so you have to validate what they say, what the expectations are, and what the model is capable of doing for them at any point in time. Another area I'd add is, especially in financial services and crucially to the bond GPT application, is that it's all about the data. When chat GPT first came out, we were amazed that you could ask it to write a poem about your pet dog in space and it would do something awesome. But for financial services, the data has to be accurate, has to be real time, and it has to be And not just real time, but it has to be the data that the financial services firm thinks is the truth. And so a lot of what we did was actually bring it back to a curated data source that could be independently verified and using that as a source of what gets generated by bond GPT. Sorry, the last element is actually around compliance, right? When it comes to financial services, a lot of the communication that happens between a broker dealer and a client is, you know, there's specific regulations around what can be said, what can't be said. And so for Bond GPT, one of the things that we had to do was to work with and understand what our current compliance obligations were and actually codify that using AI to have a compliance layer that reviews everything that happens coming out of the model. So those are some of the things that we did to get going here.
Roger Burkart: Well, that's a super set of practical experience and with a very compressed timeframe. I recall that this got to market in under two months. So a lot of learning in a short period of time. Maybe I can turn to you, Brant, and ask you to pick up on the points that Joseph was making about how to incorporate a firm's own data, right? So BondGPT incorporates bond data, incorporates some of that is open and public, comes out the filings or purchase data from companies. market data providers some of its proprietary um so how looking more broadly what are you seeing uh our approaches that firms are taking to incorporate their own data into the outputs they provide back to customers or potentially even to train the models fine fine tune the models
Brent Swidler: yeah so um there's two generally two different approaches here that um at a high level that we see companies taking across the board and my perspective is not just from the point from financial services but across different uh industries and applications here the first is through prompt engineering in the way that you um you've done this in bond gpt which is through a functionality called rag which is retrieval augmented generation which is when you receive a prompt, the model decides that it needs external information, and then it goes and executes a search query upon your proprietary dataset. The search results come back into the model, and then the model reads all of the search results and gives you the correct answer. It's what you might see with doing a web search or if you've used Bing with the ChatGPT where it's going out and doing a web search and coming back and reading the results to you. That is the main way that we see a lot of companies interacting with the data sources, because that's a very lightweight way of combining your proprietary data set with a large language model. And the way that Joseph mentioned it is that you can confine the model's response to only the data that's received from your proprietary data set. So you have a significant reduction in the amount of what we call parametric knowledge of the model being injected into the actual response. Parametric knowledge being that if I just asked it a blanket question, like, you know, should I invest in this stock? That sort of thing would be the model's going to try and use its own understanding of the world in order to answer that question. Whereas if I, you know, gave it a bunch of information and I said, is this a good investment? It's mostly leveraging the information that's provided there.
Roger Burkart: Maybe I can sort of pick up on that because we've had a couple of audience questions. One of them was around regulatory interaction to get bond GPT to market. And I'll just ask that very quickly to say I know that Joseph had to work with the firm's compliance officers to make sure that it couldn't respond inappropriately. And so the example you just gave, which is, is it a good recommendation? What bond would you recommend? We had to use an LLM to actually say, no, I can't give that kind of advice. Right. Yeah. Now, in the case of Bond GPT, as we heard from Joseph, it's using data that's curated, known to be true. It's not going out through an LLM to effectively use data from the broader Internet. But one of the audience questions was, what about copyright infringement? Is there a risk, a legal risk in terms of copyright in using a large language model? And Martin, I'm sure this will be a discussion item at ATLF US and with your clients. Can you comment on that?
Brent Swidler: Oh, is that to me?
Roger Burkart: Yes, I'm giving you the softball question.
Brent Swidler: I thought you said to Martin. Yeah, so... Sorry, sorry. The... I'm not in a position to comment on any of the legal risk affiliated with copyright infringement at all across the board. So I'm going to have to punt on that one specifically just for my own employment sake.
Roger Burkart: So I should say, well done from the audience for coming up with the kind of tough questions that we can't necessarily opine. So talk to your general counsel, I guess, is where we're going to go with that.
Brent Swidler: But I would say that I can point you to documentation as the best thing I can do for Bedrock across the board to say that there's well-documented FAQs along the risk, just from the AWS side, on the different perspectives on... copyright information coming from models. Because we also have different providers in place, right? So every provider has different training corpuses and all of that information. So there's general information on Bedrock in the documentation.
Roger Burkart: Superb. So we spent a lot of time talking about the things that we can do with generative AI, all the opportunities and some of the tradecraft and getting through concerns about risk and compliance and data and access and so on. But more fundamentally, what are the limitations of generative AI? What showcases should you steer clear of or approach with specific caution? And I think I'll start with Joseph on this one and then turn to you, Brad.
Joseph Lowe: Certainly when we think about generative AI and and LMS in general, you know, we have to recognize what they are. There are large language models and they're They're essentially predicting the next word, predicting the next phrase, next token based on the probabilistic nature of the context provided, so on and so forth. So they tend to be good at language, but they tend to not be good at like math. And they tend to not be good at specific data analysis without specific programming around it. So sometimes I hear people say, hey, can we use a large language model to predict the uptime of my application? The answer is not really. Or it's not reliable. It's not trustworthy. So I think it's important to understand, first of all, is the use case a large language model use case? Is it a language-based thing? You know, versus playing to some of its strengths. You know, if it's language-based, it's really good at summarization, you know, elaborating on something, you know, understanding free text, understanding freeform stuff. That stuff is super good, but it's not going to be good at, like, math-based things, where sometimes a basic linear regression, a gradient XG boost, that stuff is going to work much better, right? And second thing, I guess, is also being really cautious about, again, back to regulated industries, is actually making sure that you understand the scope of the data, whether it's coming from the model, whether it's coming from your RAG, and what that can come out, what that can... You know, and how that affects your user. Users love to trick, you know, LLMs, like to like, you know, try to do jailbreaks, love to make it say bad stuff. And when an enterprise, when a financial services firm is releasing this stuff, there's a lot of gotchas there to watch out for.
Roger Burkart: Yeah, it's a problem with kind of any situation where you're using it to interact with customers, right, really in any industry. I just want to emphasize the point that you made there, Joseph, which is that generative AI doesn't replace all the other kinds of AI, right? So for many years, we've been able to use AI to predict things. You know, the most basic machine learning model is that linear regression that most people got to do at college, right? But all the way through to advanced deep learning neural networks, there's still absolutely a case for models that are trained to predict particular business outputs or indeed to optimize, to be prescriptive, right? So I think it's a tendency in the technology world when something new comes along for people to oversell it and say it does everything, right? Slices and dices and bakes and all the rest of it. When, in fact, it's got some specific purposes. So I think one of the things that Joseph's saying here is you don't use it for predictive analytics. Right. We've got a whole tool set for that. And that doesn't go away. Let me turn to you, Brandt. Where would you suggest some caution?
Brent Swidler: Well, the most common limitations that we always talk about are first, hallucinations. This is where the model injects new information into a prompt. Like, for example, if I say write an email to my boss telling them something and it creates a name for my boss or it injects a date and time for a meeting or something like that, those sorts of things are limitations of models where they'll end up hallucinating content that you might not want to be there, which is why it's always recommended to have some sort of human in the loop to make sure that the content is valid. I mean, there's been scenarios in the world where lawyers have used this to write legal briefs, and it made up previous cases that were there. So there's scenarios that you'd want to reduce the hallucinations. Factual inaccuracies as well, especially along the lines of areas where there's not that much public information about a specific topic you'll end up seeing models not be the most factually accurate that's why a lot of times they want to be connected to external data sources because if you ask them about current events or you know what day is it today a model is not going to know that information because it's generally static unless it's able to leverage information from an external data source One of the things that I spend a lot of time with my team on is prompt composition. So if you say write a blog about X, then it will write a blog. But if you say create a blog or imagine a blog about something, just those slight tweaks in words can have drastic impact on what the model thinks its intent is to end up doing. So in certain scenarios, it could be Just changing one of those words could make it more imaginative versus more direct in the factual information that it ends up to generate. One thing that I did internally was for me, I said write an e-mail to my girlfriend telling her that I'm excited to go on vacation with her and the model created email, but it said that we were going on a family vacation with my parents and her parents. I changed it to a room, I just said romantic vacation. And then the new output was something that is a lot more of along the lines of what I wanted. I didn't give it any guidelines on like, what kind of vacation we're going on, or, you know, what we were going to do. But it had to create that information, because I didn't give it guidance on exactly what it is that I wanted it to do. How did she take the email? Go ahead.
Joseph Lowe: How did she take the email?
Brent Swidler: Loved it. The second one, not the first one.
Roger Burkart: So I think that emphasizes a point that is worth emphasizing, which is one area to avoid using gender of AI is a sort of fire and forget. I mean, generative AI, when we talk about content generation, we should think about that as creating a first draft. And we need to have a business process so someone reviews that first draft before sending it out. And that's what you just went through in a kind of personal example with the email to your girlfriend, right? So thinking through who's going to review this content, recognizing it is draft, and what are the consequences if it's not exactly right? So I think that's just a very important thing to be thoughtful about when you're using a large language model to generate content. And what we heard from Joseph before on Bond GPT, where, among the other things, he's using large language models to understand language, is some approaches that he's taken to make sure that the model has understood the intent of the trader and played it back so that we make sure that we didn't misunderstand it. So in both cases, the review cycle, I think, is really, really important. Good. Let's move on and talk about some practical advice for financial services firms that are looking to get started. You know, we have a third of the audience who haven't really got started yet, judging by their own hands-on experience. So what are the three things that financial services firms should be doing to get started? And how should they organize to get started? So let's start with Brent and then turn to Joseph, and then I'll fill in anything else that I think you can help with.
Brent Swidler: So I may have mentioned this before, but the first thing that we always we see a lot of companies doing is just starting out with an open ended internal playground that people can experiment with different. You know, you can democratize access to all the different all your employees to drum up different use cases. And not only do they get their own productivity improvements, but they often are creative to figure out exactly what is they need to do in a certain day. and create a list of valuable use cases that each employee would end up seeing as something that they internally would be able to benefit from. That's usually the first step that we end up seeing across the board, and that's something that we do internally as well. The next thing is that with the use cases that are developed, they should funnel into some development teams that would be able to productize or implement some of these applications across the board for different folks still remaining internal to a business. And then once you've looked at the spectrum of different use cases, you can, you know, we talked about this with the example of amazon.com summaries, but you can look at them at a, you know, easy to implement low risk applications versus more difficult, challenging applications that require a lot more interactions from an end user and be able to prioritize where you develop along that spectrum. And then the third thing is along the lines of change management, at least educating your – educating your users on exactly how they should be interacting with these tools. So I'm sure that in the case of, you know, you can correct me if I'm wrong, but in the case of Bond GPT, you don't want people making decisions without having any interaction, you know, having any review of the output that's generated from the models themselves. So just like the third thing is definitely an education path and making sure that people are using these things correctly and, you know, reviewing the output that they're generating, whichever application that they're on.
Roger Burkart: Super. Thank you, Brent. Joseph, what's your take on what were your top three?
Joseph Lowe: Yeah, you know, it's similar to what you were saying, Brent. I mean, the first thing is, you know, you call it democratization. I'd call it like equipping your users. The first thing is ensuring that your users understand the capabilities of this new technology. At Broadridge, for example, we launched an internal playground, we'll call it. We call it creatively broad GPT that has access to open AI models, has access to LAMA2 models, and coming soon, AWS Bedrock models, where users can use it for productivity, they can use it for testing out use cases, do that all within an environment that is secured and within our Broadridge guidelines. But what that does is it creates the knowledge that the team needs in terms of expectations, in terms of what's possible, right? I think that's really important because a lot of the things that we're, you know, with the generative AI, we're seeing that it's not going to be just the people in the AI tower that know what to do. It's actually every single business line, every single product area, you know, people are experiencing AI and using that. So for us, after that, the second thing that people need to do is help their teams get educated, not just on what the AI can be good for, but how AI is used in delivering products to customers. That's where we have to think about things like evaluating model performance, understanding the costs associated and the infrastructure associated with delivering products. well-made applications that leverage these things. And that's something that, you know, Bedrock helps with when it's a managed service, right? But those things are real concerns. Considerations for IP, considerations for just product management. How do we make things better over time without, you know, how do we deal with, you know, when models change behind the scenes? So all that stuff that has to do with making a good product that lasts. And the third thing I'd say is it's really important to get your internal stakeholders involved. With all this technology, what we've seen is that it's just that technology is going to be fine, but we need to make sure that especially financial services, that people are involving their legal teams early on, their compliance teams, their risk teams, all the different stakeholders when it comes to interfacing with your customers, interfacing even with internal data. all that stuff is really going to matter. And what you don't want to do is have a great prototype, a great POC, a great pilot, but unable to, you know, deal with some of those issues.
Brent Swidler: Go ahead, Brad. Sorry, go ahead. I was going to say, I was going to add to one of the things you just mentioned, which is on model evaluation, which is something we go into a lot of depth with across the different companies that we work with, especially being Bedrock, where we have a wide variety of different models provided inside of the Bedrock service. When we do our own internal evaluation, there's definitely three different approaches that end up happening here. One is that there's a lot of external benchmarks, which are usually automated ways of evaluating models, and I'll touch on that in a second. We develop our own automation for model evaluation, but I rely most heavily on our human evaluation of model outputs. The reason is because a lot of a lot of this information is a lot of the output is highly subjective. And when you do automated evaluation, what you're normally doing is you're looking at the output and trying to judge the output in some sort of automated fashion. Like, is it saying yes or no? Is it saying true or false? And even in those scenarios, we're seeing a lot of instances in which it's really hard to create automated evaluation of specific models. Like if I ask, is California bigger than New York? The first thing you'd say is, okay, I'll just look for the word yes. And so if the model outputs the word yes, then it gets it correct. Then the next model will say, yes, California is bigger than New York. And in that scenario, yeah, I got it correct. But you might, depending on how you're splitting up the content, you might say that that's now incorrect because it elaborates further. So you go to a partial match. And then the next one says, yes, I like LAMP. So it technically said yes, but then it went off the rails afterwards. And so did that get it correct? And then what we see a lot of models end up doing is say, California is bigger than New York. So it didn't say the word yes. And so a lot of the automated evaluations or the different frameworks people use to evaluate models become broken in that nature, which is why we end up evaluating models on the human evaluation more than just the automated metrics.
Roger Burkart: Super. I usually like that example. One of the themes for both of you is the importance of education, right? And in a former role, I was partner at McKinsey. We did research on the impact of AI and which organizations were effective with AI. And what we found was that a sample size of about 1,000, it's only about 14%, really felt they'd been successful and were showing up in their financials. And the rate limiting factor in many cases was the sort of executive education, the business education to enable business leaders to understand how can I get value out of this? Either as someone who's building products or someone who's running an operation, someone who's trying to get sales done, marketing done. And so investing in your business leaders and your product leaders is really important. And the nice thing about generative AI is in many ways, it's somewhat lowered the barrier to entry technically. So we're not having the same kind of conversations about how many unicorn data scientists do we have to find to build predictive models. But it was still the case, I think, that the rate limiting factor for many organizations is getting business leaders up the curve. and giving them opportunities to learn safely, as I think you started with Brad. And I'd also just add a plus one to a point that Joseph made about the importance of collaborating with your stakeholders. We need to recognize if you're in a financial service firm or other regulated firm like the health industry, that this is a learning journey for the folks in the control functions too, right? Just as this is new for technologists, it's new for business people, it's new for the legal fraternity, it's new for risk management or model risk management. And so working out how can you start with relatively low risk I think that was one of your advice items earlier, Brad. How do you choose some early ones that aren't too risky, aren't perceived as being risky? And that helps you bring your control functions up the learning curve. And they build confidence that they know how to manage this and move forward. Two other things on the education front. What I'm seeing happen is that users are starting to recognize sort of safe patterns of using large language models. And once you've found a safe pattern that the organization is comfortable with in one business area, you find you can apply it in many other business areas. So Joseph's pioneering work with Bond GPT is now being adopted at Broadridge in other trading applications, in operations applications, in regulatory reporting applications. And it basically recommends, hey, there's a pattern here that is safe for us to use, whereby we're using a large language model to provide a safe way of navigating inside data, which we know to be correct, right? So if you can start with things that you recognize are safe and reproduce and share those patterns, that is a very, very important way of moving forward. And finally, on this education point, it's important to recognize what's new and what's new. And so I think Brant and Joseph, you both talked about prompt engineering and what you can do with prompt engineering. This is a new way of working with AI, right? Prompt engineering is seductively straightforward, but there's also a lot to learn to get the most out of it. And excuse me, Brent, you mentioned a couple of examples of that earlier on in the nuances of how you instruct the model. I think over time, what we'll see is organizations will have libraries of prompts, templates that are used to make sure that individuals get a value out of the way they're interacting with users. So we're going to move now to another poll. And then while that is calculating and while you're providing your input to it, we'll then move on to another question. But if you can just sort of give us your responses to this, it'll be very interesting to see where this audience is at in terms of having plans, pilots, or indeed production use cases. So while people are working on that, we have a kind of structural question for Brant and then Joseph, which is, as generative AI has democratized the access to powerful AI capabilities, so we no longer need large teams of data scientists and long projects with big budgets to get a predictive application out of the door, do you think it will level the playing field? between the folks with the big IT budgets, and AWS is one of the largest, and smaller firms who don't have such deep pockets. What's your sense, Brent?
Brent Swidler: Well, you gave me a good seed there because in the one case where there's scenarios in which you want to become a model provider, yes, you're going to have to remain having a deep pocket. because you see all of the different companies raising billions of dollars in order to execute this space. So it's very rare that there's going to be companies that want to actually get into the business of being model providers, as in here's a model that I want you to use and leverage and sell. Now, on the other end of the spectrum, we're talking about just applications in which – you know, that you're using existing models out of the box, it absolutely is democratizing the ability and is leveling the playing field to some extent. Now, there's still applications that need to be built around these models. So it's hard to necessarily say that, you know, there's like folks with bigger budgets are not going to be able to execute faster and more efficiently and in a more robust way than ones with smaller budgets. But It really just depends on the application and the use cases that you're going after and the way that the business transforms around it.
Roger Burkart: Joseph, what's your sense? Is this going to let the little guys compete more effectively?
Joseph Lowe: Well, I kind of agree with Brett more. I think there's a difference between creating new foundation models versus usage of AI, you know? And I'd say that, you know, when you look at the costs of, you look at the costs of, you know, opening AI GPT 3.5 or Claude, you're talking about 0.2 cents for a thousand tokens, you know? And I think that democratizes it in terms of anyone that wants to use a large language model, cost is not going to be a barrier at the lower end of the scale, you know? I think people have to think about when things do scale, when you do have all your users using it, what does that look like? I think for sure getting started has been democratized. Again, it doesn't replace traditional data science. There's a lot of need for that stuff, and that stuff is still going to be very much a data game, very much a skills game. Actually, that's probably a good point is that Yes, Gen AI democratizes the tools, but the firms with the big data and who have done a good job collecting their data, organizing it, structuring it, will have an advantage. And the firms that have been doing this for a while are going to have an advantage there. But I think Gen AI helps to level the playing field.
Roger Burkart: Super. Super. Well, let's hope for the smaller guys here. It's good to hear. So let's look at the results from the query. So we have, you know, only 8.2% people have no plans to use this. So this is an action-oriented group here. We do have a group who prohibited the use. And I wonder if that has to do with some of the regulatory and compliance concerns that we have. So some work to do instead of coming up with safe ways to roll out tools. And then, you know, 31% and 28% either planning or planning to pilot or deploy. So 17% with actual multiple live use cases. You know, that's pretty much what I've seen. We did a poll at the Futures and Options Conference a couple of weeks back, and the I think we had 10% there. So this looks pretty in line with what I've seen. But what about you, Brad? You cover many different industries. What do you see?
Brent Swidler: I think this jives well. This seems to be a good bell curve of folks still in the planning stage and going through different pilots. And then a smaller percentage of the ones that have no application and a small percentage of the ones that have live applications. I think in the next year, we'll start to see these numbers skew more towards the live applications, obviously. But this seems to be along the lines of what I would expect. The interesting thing is that I'm trying to overlap the previous poll with this one. which is that there's 30% of people that don't use any generative AI tools, but that seems to be different in the actual organization, which is in different stages of planning, which is interesting to see.
Roger Burkart: Super. Joseph, quickly, any quick add-ons to that?
Joseph Lowe: Yeah, I mean, the one thing I'd add is that more and more applications that businesses are going to be using from vendors are going to have incorporated Gen AI. So I'd say that this number is probably skewed in that, you know, people that are using Office 365, you know, that stuff is already seeding into there. And so I think that we're going to see that it's not just going to be, hey, I'm doing a Gen AI project at Business X. It's going to be products are going to be released there, and it's going to be very commonplace.
Roger Burkart: Yeah, we often underestimate how much we already use, right? You can have a conversation with someone at AI and say, oh, I'm never going to use AI. And they turn around and talk to Siri and ask you to order them a hamburger.
Joseph Lowe: Exactly. So, for example, for all the firms that we're talking, that we're working with on bond GPT, they may not be considering that a Gen AI project. They may be considering that as a insights tool, pre-trade insights tool for their bond traders, right? And so I think that's to be considered.
Roger Burkart: Super. Well, let's turn a little bit now back to the case of how we use this, what value we get out of it. And I'm going to ask Joseph and then Brand to talk a little bit about, do you see generative AI as being more of an efficiency driver? Or do you see at some point that will actually be transformative? In other words, that the impact would be big enough or could be big enough that firms will need to reinvent their business models?
Joseph Lowe: I'm going to play a little devil's advocate here, I guess, if I can say that, is that I think Gen AI is going to be the impetus for the firms to want to get their digitization programs in place. I think when you see generative AI, which to me really is... it allows for a change in user expectations. It's a different way for users to interact with the applications, the data, the services, the workflows that are there. What that really means is that underneath that user experience layer, there's going to need to be a lot more digitization and transformation. All the stuff that firms are struggling with that are working through today, Gen AI is just going to demand acceleration of that in order to really capitalize on some of the benefits that we're seeing.
Roger Burkart: So Brent, what's your view on this one?
Brent Swidler: So there's definitely businesses that are being massively disrupted by this new paradigm. what we see across the board from different companies is and it might just be the stage of where the products and you know the rollouts are as we see 17 of people only have you know companies have live applications but it's not necessarily the business models that are being transformed just yet but the products and user experiences So just some examples, if you just because these are all public, if you go to Amazon dot com, if you're a seller, there's services now where you can like have the you can have uh models generate product descriptions for you or create your listing images and all of these different things now make it easier for sellers to you know create optimized listings so it's not that there's a change in the overall business model but the products that we're seeing across like a number of different use a number of different businesses and amazon's just you know the amazon.com is just one of them is um making it so it's a completely different user experience and uh The applications are a lot easier for end users to be able to use, and now it makes creating a listing a lot easier. Across the different companies that we work with, we end up seeing generative AI use cases being implemented into products that simplify applications for users, whether it's searching for different things or generating images in certain scenarios. It's not necessarily the business models just yet, but I imagine that that's probably to come next.
Roger Burkart: Right. And just one thought on this is that I think many organizations are starting with efficiency use cases because they want to start in an area which is lower risks of internal use rather than things that get to the very fundamental relationship with their customers, right, which may have more transformation. I was part of a rather fun debate among senior wealth management executives about whether this was really a productivity tool, efficiency play, so to make financial advisors super effective, providing personalized content and then advice to their clients, or whether it was truly transformational. And what I came away from that discussion, because there were strong voices on either side, is that At some point, when the productivity gains become enough, you actually are changing the business, right? The wealth management space is a big battle between the more robo-advisors and the more human-led advisors. And if the human-led advisors are truly much more productive, then at some point it tips over to becoming transformational. So I think we're going to start to sort of focus a little bit more on some of the user questions here. One question that came in early on was a little bit around how should we think about the concerns about bias with generative AI and perhaps indeed more broadly in AI? Let me start perhaps with you, Brent, and then Joseph, and then perhaps I'll share a little bit of my own experience with trying to avoid bias in AI models.
Brent Swidler: I guess I would want to ask bias in what content, in what context, because there's biases towards like, you know, there's bad biases in which you say, like, you know, write an email to this doctor and the doctor, you assume the agenda of the doctor in certain scenarios. But then there's biases and in other scenarios where like, It might opine on certain topics that you might not want it to and have a perspective or be skewed towards one political affiliation or another. The way that we see a lot of models ending up working is that you can somewhat manipulate the bias through the way that you engineer the prompt. And if you're trying to reduce the bias, you can actually, you know, you can fix the results the way that the model generates content to the data that you've provided it, as well as the instructions and system prompts that you end up giving the model to, you know, confine its response structure. Now, the models predicting the next token are based on the data that they were trained on and which, in a lot of scenarios, is publicly available information. And so it's somewhat of a reflection of the content that the world has produced over the number of years. And we just have to manipulate the models through the way that we engineer the prompts to generate the right result. All the model providers will probably train to make sure that it reduces bias as much as possible, but there's always ways to elicit or control for the biases as much as possible.
Roger Burkart: Right. And maybe I should have taken the audience question and sharpened it a bit because I think what we mostly care about is really harmful bias, right? And so the context, as you said, Brad, is really important. In the world of predictive analytics, predicting whether someone is likely to repay a loan is a classic case where you could get harmful bias because the models are trained on data of human behavior, and that could be a real problem. And so in the U.S., we've had for decades now the Fair Credit Reporting Act, which really requires banks to be very, very careful in proving that their model works. doesn't have a higher probability of giving a loan to one ethnic group versus another, for example. So that's like a high stakes area where in predictive AI, where we've got quite a lot of experience. And one of the things we do there is we're really careful about what data we use. And we recognize that the data does have embedded in it some biases and we need to understand them and often correct for them. And I think one of the things that's very different about generative AI is we are rarely getting the chance to choose the data that trains the model. And so maybe we're using it in a lower stakes environment, so it's not going to lead to someone not getting a mortgage, not being able to buy a house, for example. But nevertheless, we don't have some of the same control that we're used to having when we choose the data to train a machine learning model, let's say, to make a long decision. Joseph, any thoughts on your side on the bias question? Or should we turn to user experience, which is an area about a lot of questions.
Joseph Lowe: I would just say that, yeah, the bias is going to come up in the data that is used to train the models, as well as the RLHF, the human feedback that goes into aligning the models. Bias can be introduced there. And I think anyone using these models needs to take that into account, understand what's good and bad, and go from there.
Roger Burkart: Super. So we've got a number of audience questions around now. How do we go about thinking about the user experience or the UX? How do we interact with users in this world of chat with large language models? And I know this has been a rapid learning curve within your world, Joseph. And some of the things that you've come up with, I know, have been widely adopted elsewhere. Maybe you could talk about a little bit of how you went from having a single little box saying, ask me anything about BOMs, to the user interface you landed on with BOM GPT. What are some of the key learnings about how to build a user interface that is safe and effective?
Joseph Lowe: Yeah. Certainly, chat is just one application of large language models, but I'll talk about the chat interface. For us, it really comes down to aligning the user's expectations. When a user asks a question and they want to know something, first thing to do is really make sure that we can replay back to the user what it is that we understood them to be asking about. I think that's kind of check one. Did we understand you? Number two, with BonGBT+, which we announced last week, it's about explaining to the user what steps we took to get that answer for the user. Third is being really clear when you display the answer, where the sources come from, attribution, but also from there being able to provide feedback to the model or to us around potential inaccuracies, potential misunderstandings of the question. For example, if I said, you know, If you ask the question, you could take a meaning to be both ways. We make sure that it's obvious to the user which way the model interpreted to be, where the answers are coming from. Last thing that, you know, we really kind of focused on was around, you know, did you mean? So when we think we might have misinterpreted the user's question or it could have been A or B, like maybe a 60% A, 40% B, you know, what we'll do is we'll actually say, well, did you mean B? Did you mean C? You know, and that helps the user understand that in a chat-based interface where even the inputs are not deterministic, that we give the user an out for actually correcting the model.
Roger Burkart: That's really super. I mean, it is just such a different world where you have such a powerful tool that allows you to basically guide conversations. I mean, we're used to user interfaces where you click on something, you choose a dropdown, and it's like a transaction, right? So you're kind of designing paths, proper paths that a conversation might take, which I think is a really new and different design question. Brent, any sort of final thoughts from you on the user interface design? Yeah.
Brent Swidler: So my user is different, in which I have to design a product for someone like you all to be able to define the guardrails and applications that you want to build and the flexibility that you want to give your end user. So I make sure that our models and our service are functional to you all to make sure that you can define the breadth of the different applications that you want to be able to end up having, which means that the user experience that we have, whether it's chat or just a base text box, is built with a lot more flexibility for an end user, which, depending on who you guys are interacting with, you might want to reduce that scope significantly to only have conversations about bonds, or you might want to have a much more broader conversation about a bunch of other different things or link to web searches and connect to other stuff. So we've built for a lot of flexibility so that you can end up building the application for your end users.
Roger Burkart: Well, thank you. And, you know, our time is up. This has been a really interesting conversation. Thank you all for joining. I hope you found this useful. Many of you have plans to roll out applications soon, and I hope some of these practical learnings will be helpful to you. I think we're all on a learning journey here and will continue to share through various forums. Let me close by really thanking Brent and Joseph for sharing their expertise so clearly with such practical examples of how to be successful in this field. So thank you all. Much appreciate you being part of this conversation.
Brent Swidler: Thank you. Thank you.