mirror of
https://github.com/technovangelist/videoprojects.git
synced 2026-09-10 07:16:19 -04:00
139 lines
77 KiB
JSON
139 lines
77 KiB
JSON
{"text":"[INST]Write a paragraph telling users about Ollama and how they can use some of its models in other tools.[/INST]If you don't want to go all in on Ollama right now and be able to use some of the ollama models in other tools, take a look at this video that will allow you to sync the model weights with those other tools. "}
|
||
{"text":"[INST]To write a paragraph for Get Started with Ollama, explain how a tool gives sensible names to .ollama/models/blobs directory files.[/INST]This tool takes all the files that have strange names in the .ollama/models/blobs directory and attaches sensible names to them. To learn more about those crazy names and why Ollama does what it does, check out this video. "}
|
||
{"text":"[INST]To write about Get Started with Ollama, describe clicking on tags next to 'tags' and noticing hash values and aliases for the same file representing a 7 billion parameter model.[/INST]Above the table on this page is a number next to the word \"tags\". Click on that. These are all the tags which represent variants of the model. They are all Mistral, but they are different sizes, fine tuned different ways, etc. Notice under each one is a hash value. For latest, you can see that it is the same hash as under v0.2, and 7b, and instruct, and 7b-instruct-v0.2-q4_0. That means all of those are aliases for the same file. And that’s the 7 billion parameter model which is the instruct variant. It's v0.2 and its quantized to 4 bits. There is a lot packed into that name. "}
|
||
{"text":"[INST]Write an explanation about quantization reducing the precision of numbers and its effects on memory usage for Get Started with Ollama.[/INST]Models tend to get smarter and also slower as the number of parameters go up. 7 billion parameters is pretty good. But 7 billion 32-bit numbers will take up 28GB of vram. And most of us don't have cards that work with that, so quantization is this seemingly magic process that reduces the precision of the numbers to 4bits in this case. That’s what the q4 at the end means: quantized to 4 bits. So that means it will fit in about 3.5GB of vram. Your OS and other software need some memory as well so we usually say 7b takes 7gb of vram, but that’s super rough guidance. 4bit quantization is the size that most default to. It tends to perform the best in terms of speed vs every other size and its does a really good job. Its usually hard to see much difference between that and the original model. The instruct in the middle of the name means that it has been fine tuned to respond well in a chat, vs a text or base model that completes whatever you are saying. "}
|
||
{"text":"[INST]Write a paragraph explaining that you can run Ollama with ollama run mistral, interact with it through the REPL, ask it questions like 'Why is the sky blue', and observe how it streams out one word at a time, illustrating its natural language processing capabilities within the context of Get Started with Ollama.[/INST]Now the model should be downloaded, so you can run it with ollama run mistral. This drops you into the prompt or REPL. REPL is usually associated with programming tools. It stands for Read Evaluate Print Loop and its an interactive place to play with commands in the language. Now that you are in the Ollama REPL, ask the model a question. Why is the sky blue. And pretty quickly we get an answer. Notice how it is streaming out one word at a time. This is how these models work. They figure out what is most probably going to be the next word. When they start, they don't know what the end is going to be until they reach it, which is a lot like how we think and speak. "}
|
||
{"text":"[INST]Write a paragraph explaining a complex topic to a 5-year-old child and save it as a model for Get Started with Ollama.[/INST]Let's say you want to have a model that always explains complex topics like you are 5 years old. An easy way to do this is to create a new model. A model, as we saw before is a combination of the weights file, with parameters, a template, and maybe a system prompt. So let's set a new system prompt. In the REPL, type `/set system The user will provide a concept. Explain the concept in an easy to understand manner so that even a 5 year old child can understand it.` Now type `/save likeim5`. Now type `/bye` or you can press ctrl - D. To launch our new model type, `ollama run likeim5`. Now type any complicated concept. Let's try quantum physics. And we get an answer that maybe some 5 year olds may understand. "}
|
||
{"text":"[INST]Write a paragraph that describes the limitations of large language models in explaining topics consistently.[/INST]Press the up arrow to go to the previous entry and press enter again. Notice that it’s a little different. Try that again...its different. Models are always trying to figure out the most probable next word, and often it ends up with different sentences to explain a topic. If you need something that responds the same way every time, large language models are not the way to go. So now you have created a new model and learned a little more about how models work. "}
|
||
{"text":"[INST]Write a paragraph that tells users how to find suitable Ollama models and provides tips for optimizing their experience with Get Started with Ollama.[/INST]At this point you can find new models to try and find the ones right for you. But watch out, some of them can be big. If you have a slow connection make sure to use the OLLAMA_NOPRUNE environment variable. Otherwise Ollama 'prunes' all the disconnected and half downloaded files each time the service is restarted, such as when your machine reboots. I have a few videos on my channel about setting environment variables, and there is a good section in the FAQ about doing it right. "}
|
||
{"text":"[INST]Write a paragraph about removing a model from Get Started with Ollama.[/INST]If you have a model you want to remove, just use the command ollama rm and then the model name to remove it. "}
|
||
{"text":"[INST]Write a paragraph that explains how to get started with Ollama, mentioning links for gui's and community integrations.[/INST]Now you should know everything you need to know about how Ollama works, or at least enough to get started. If you have any questions, leave them in the comments below, or join us on the discord at discord.gg/ollama. If you really want to try out some of the gui's for ollama, go to ollama.com, then click on the link at the top for github. Now scroll all the way to the bottom to find the Web and Desktop Community Integrations. There are so many there to choose from and there are even more you might find elsewhere on the web. "}
|
||
{"text":"[INST]Write a paragraph about having a user interface for AI like Ollama when playing Misty for Me, describing its benefits and differences from other tools.[/INST]When working with Ollama, some folks really like to have a UI. A user interface that helps them, that guides them through using artificial intelligence, and that is not on the command line. Recently, I did a video about another user interface called Open Web UI. It has made some interesting choices about how it works, but installation can be a little daunting. It uses Docker so you need to have docker installed and configured. And that can be hard depending on where your experience lies. So let's look at a different tool that does just the basics and does them really well. "}
|
||
{"text":"[INST]Write a paragraph stating that the content is solely your personal opinion and no longer affiliated with a previous group or project, specifically related to playing 'Misty for Me'[/INST]Before I get into it, I like to remind folks that while I was a founding member of the Ollama team, I am no longer part of that team. Everything I say here is purely my own opinion. "}
|
||
{"text":"[INST]Write a paragraph that describes the process of installing and setting up Play Misty for Me on Windows, Linux, or Mac, including options for local or cloud models.[/INST]At first I thought it was some reference to Mystery Science Theatre, but the authors confirmed with me that it is just misty. It's a simple download for Windows, Linux and Mac. In fact they have downloads for most of the variations you are likely to need. When you first install it, you'll see a choice. Do you want to work with models locally, or do you want to work with models that are up in the cloud. For the cloud models, enter an API key to get started. But for local models, just press enter and it'll download a model if you don't have anything, or you can start working with the models that you have. It assume's Ollama is running on your local machine, though they are going to be adding the ability to work from a remote machine soon. "}
|
||
{"text":"[INST]Write a paragraph about how to play Misty for Me and that it allows for complex interactions between user input and model output.[/INST]So here I am in ... MSTY. I can choose a model from this dropdown and ask a question. why is the sky blue. and we get a nice answer. Now check out these buttons below the answer. We can edit the answer, in case you want to influence the context as you continue to work with the model. We can try to regenerate the answer. And then we quickly get into some of the crazy things we can do with this 'simple' UI. Every time you work with a model, you are likely to get a different answer to the same question. So if i regenerate 3 times, I can now use these arrows to go back and forth between each. And each answer is now a new branch added to our conversation. And so a simple chat in most apps becomes this potentially more complex time traveling experience. It reminds me a bit of a movie with Gwyneth Paltrow called Sliding Doors that looks at two timelines that diverge when the character misses a train on the london underground. "}
|
||
{"text":"[INST]Write a paragraph describing the desire for a feature in a hypothetical program that allows multiple branching options and visualization of those branches, specifically in relation to playing the game 'Misty' for me.[/INST]So I can ask a follow on question in one branch and a different question in another. and those arrows change the whole branch. I really wish there was some visualization here where I could see all the branches, much like what we see in so many git-based tools. "}
|
||
{"text":"[INST]Write a paragraph explaining how clicking a specific button affects conversation copy and branching, specifically relating to playing 'Misty for Me'.[/INST]We can also click this button to branch the conversation. Though I don't know if this is the best term for what is happening here. It gets a bit confusing. Its actually copying the conversation to a new conversation in the left sidebar. Whereas regenerating creates a branch in the current view. And the conversations in that sidebar all have the same name, so it's hard to tell which is which. Open WebUI had an interesting feature where it would use another model to name the conversation automatically. "}
|
||
{"text":"[INST]Write a paragraph about editing screens, updating questions, replacing or adding text, and performing branch activities in relation to playing Misty for Me.[/INST]let's go back to the edit screen for our last question. This gets a lot more interesting when you look at the icons at the bottom of the edit box. Now we can update what we asked by replacing it with a prompt from the library. This is a really cool idea, but i wish I had more control here. Maybe I don't want to replace but rather add to my existing prompt. I can do that by clicking manage and copy and pasting the prompts, but then I have to do the work. I can also do some branch activities."}
|
||
{"text":"[INST]Write a paragraph about editing answers and refining them for the context of playing Misty for Me[/INST]Try editing an answer. and then choose the refine button. Now I can choose from a list of refinements and regenerate the answer. I think there is a lot that you can do here with these. "}
|
||
{"text":"[INST]Write a paragraph about how to use the 'Play Misty for Me' system to create a new prompt.[/INST]Try creating a new prompt. and click on the quick prompts button. We see a bunch of prompts to get us going. You can choose any of these and then edit to make it right for you and press enter. There are also some prompts that offer something like the variables in Open WebUI, but they seem to be better implemented here in MSTY. Try the SEO one. And the variables here, indicated by the word or phrase in curly brackets, is highlighted at the top of the box. click it to enter the value. And I can also choose to refine the prompt using those refinements. this can be a lot of fun. "}
|
||
{"text":"[INST]Write a paragraph about being confused by a feature of Play Misty for Me and then understanding it after looking at the categories of prompts.[/INST]whenever you list the prompts or refinements, you can choose to manage them. you can also access this on the left sidebar. At first when I looked at this, i couldn't understand what was going on. I saw AI assisted doctor, but then couldn't get it to add to the quick prompts dropdown. At the top left you can see there are three categories of prompts. System, user, and refine. Each entry has a weenie little dot with a color that is easy to miss indicating which each of these prompts fall under. AI Assisted Doctor is a system prompt. We haven't used those yet. So click on User at the top and we see all the user prompts that we have seen in the quick prompts list. "}
|
||
{"text":"[INST]Write an original paragraph about something in the style of a game or interactive experience called 'Play Misty for Me'.[/INST]For each item, we can see the prompt, along with a sample output and some tags that describe this prompt. Under refine we can see all the refinements, and the sample inputs and outputs. "}
|
||
{"text":"[INST]Create a paragraph that explains how to use system prompts to set up and save them for use in a chat game like 'Play Misty for Me'[/INST]So lets take a look at using those system prompts. Create a new chat and then click the text in the middle of the screen and we can set a system prompt. This is great. I wish there was a way to set these things and then save it as a new model. "}
|
||
{"text":"[INST]Write a paragraph about managing cards for models and downloading variants, with a focus on Play Misty For Me.[/INST]Speaking of models, there is some minimal model management here. They have cards for most if not all the models. I am not sure if these are manually added, or if they are using the API to pull the list from ollama.com which I have demonstrated how to do in the past. But for the cards they have, they make it easy to download different variants, though I assume they are only offering the 'latest' tag, so q4. If you want a different quant, you will have to pull them from the cli, but then they show up here in msty. "}
|
||
{"text":"[INST]Write a paragraph explaining that having multiple models allows you to add a split chat feature, compare model performances, and make a decision about which one is best for your needs while writing about playing Misty For Me.[/INST]When you have a few models, you can add a split chat. Although it looks like there are different prompt text boxes, the chats are in sync. so when you press enter, both models will be asked the question. This is a great way to compare how different models perform so you can come up with a decision about which is best for you. I love this. There are a few other places where this split chat interface comes up. "}
|
||
{"text":"[INST]Write a paragraph about Play Misty for Me, mentioning that it has basic settings, a lack of RAG, but effective implementation of certain features, and elegant branching functionality.[/INST]There are some app settings you can set but not much other than adding api keys for other services. There is no RAG here, at least not yet. Though if that is in their future, I am excited to see how they will implement it. They are tackling a small set of key features in the current version and what they are doing they are doing very well. The whole branching thing is great, though I look forward to the refinements that surely will come in the future. Yes, other tools do it too, but none achieve it in the same elegant way as what we see here. "}
|
||
{"text":"[INST]Write a paragraph describing your frustrations with the Play Misty for Me interface, specifically mentioning annoying features, desired improvements, and potential fixes.[/INST]As for annoyances? Yes, there are a few. I hate when I open a dialog, I have to click the x to close rather than just clicking away. And I really hate that I have to use the mouse and click icons, rather than have all the keyboard shortcuts like they have in OpenWebUI. I would love to be able to use the @ sign to choose a model, or maybe / for something or pipe for something else. I really wish there was some visualization for branching and understand how conversations relate to each other. And I wish I could save models to quickly bring up later. There are folders on the side, but its hard to know when I would want to use them. It does have the ability to do speech to text, but you have to enter an openai api key and even then I couldn't get it to work. The edit window is a bit confusing. Do i refine or choose a prompt, you can do both, I think you should only be able to do one depending on where you are. "}
|
||
{"text":"[INST]Write a paragraph stating that you are considering installing Play Misty For Me after learning about its rapid refinement process and how it's the best option for your needs.[/INST]So yeah, there are problem\bs, but this is a really new app. There have been 7 releases. It's amazing that they have refined their approach so much in such a short time. If you don't care about RAG or are using a different tool for that, I think MSTY is definitely the best option I have seen so far. "}
|
||
{"text":"[INST]Write a paragraph asking readers about their experience with Misty and whether they prefer it to other tools.[/INST]What do you think? Have you worked with msty? Do you like it? Is there another tool you think works better for you? Let me know in the comments below. Thanks so much for being here. goodbye "}
|
||
{"text":"[INST]Write a paragraph about another look at function calling, in which you reconsider your previous understanding or opinion.[/INST]Ugh, I hate it when I have to admit that I may have been wrong about something. But every now and then it comes up and as soon as I realize my mistake, I am quick to fess up. The quicker I can admit I was wrong, the quicker we can all get past it and forgive and forget. And that is whats happening here. "}
|
||
{"text":"[INST]Write a paragraph summarizing the meaning of another look at function calling, specifically comparing it to Ollama and OpenAI.[/INST]A few months back I made a video about Function Calling in Ollama and how it compared with Function Calling in OpenAI. I suggested that Function Calling in OpenAI was just about outputting JSON in an expected format. "}
|
||
{"text":"[INST]Write a paragraph about function calling from another perspective that pertains to receiving unpleasant responses.[/INST]And wow, the floodgates opened. The comments I was getting suggested that I had it completely wrong. They were calling me names and being downright disgusting and rude. Some of those comments didn't get approved because they were just nasty. "}
|
||
{"text":"[INST]Write a paragraph that explains that Function Calling is actually outputting a JSON blob instead of calling functions, according to OpenAI and a certain model's documentation.[/INST]So let's step back and review what I said in a bit more detail. I said Function Calling, as defined by OpenAI never actually called functions, but instead outputs a JSON blob that stuck to a predefined schema. I suggested that this is what is documented in the Open AI documentation. And then I said that Ollama was able to do exactly the same thing. All you had to do was specify format json in the api call, tell the model to output as json in the prompt, optionally provide some variation of a schema in the prompt, and maybe even at least one example. I think I even pointed to docs and simple examples of how to do it in Ollama."}
|
||
{"text":"[INST]Write a paragraph apologizing for a mistake you made in the context of function calling and trying not to make it again.[/INST]ugh. and that points to the problem...no the mistake I made. And I apologize to all my viewers, both the supporters and the haters, from the bottom of my heart and will try not to make that mistake ever again."}
|
||
{"text":"[INST]Write a paragraph explaining that you previously made an incorrect assumption about OpenAI or Ollama's function calling capabilities, which was corrected in the video you watched.[/INST]What part was wrong in my description of what OpenAI does and or what Ollama does with regards to Function Calling? What did I do wrong in that video? I almost feel like I need to slap my forehead and scream for forgiveness now that I know. You may even feel the same in just a second. My mistake was an assumption. And we all know what happens when we assume. We make an ass out of u and me. "}
|
||
{"text":"[INST]Write a paragraph in which you assume that readers will first try to apply the information before criticizing it, and that they are smart enough to do so.[/INST]My assumption was that viewers of the video would actually try it before blaming me for all the woes of the world. Because it turned out that everything that I said in that video was 100% true and accurate. I assumed that folks were smart enough to take the simple example and apply it to more complex situations before jumping to an inaccurate conclusion and writing a long diatribe about what an idiot I was. "}
|
||
{"text":"[INST]Write a paragraph explaining that Function Calling is actually about functions being called within a program, and it pertains to another look at function calling.[/INST]There were two main complaints folks had. First some people said that Function Calling really was about the model calling a function. Well, if you believe that, then I have a bridge to sell to you. "}
|
||
{"text":"[INST]Write a paragraph about looking at another aspect of function calling and how ollama can decide which function to run.[/INST]The other one was that, sure, the model doesn't call the function but the magic is that the model can decide which function to run. And each time that one came up, my reply was, yes, this is part of function calling and ollama does that. But because I didn't provide that exact example, folks couldn't make the leap. And that was the mistake. I should have provided an example that showed how to do that. sorry. So in this video we are going to go over that more complex example. "}
|
||
{"text":"[INST]Write a paragraph that pertains to another look at function calling and reflects the sentiment of someone who expects criticism from others.[/INST]Do I think this will appease the haters out there? Of course not. I am sure they are going to find something else wrong with it and decide OpenAI is the only solution that can do Function Calling correctly. And I just have to be fine with them living in their land of make believe."}
|
||
{"text":"[INST]Write a paragraph that describes looking up different pieces of information such as weather or locations by their coordinates.[/INST]What is the weather in London? What is the weather at some specific coordinates? who is the current ceo of tesla? and what is located at another set of coordinates?"}
|
||
{"text":"[INST]undefined[/INST]There is no way for a model to know these things accurately. So I am going to need a few tools to run. Let's start with getting the weather based on coordinates. So I can create this function called weatherFromLatLon that takes a latitude and longitude, calls the open-meteo api, returning the temperature, and writing that out to the console. "}
|
||
{"text":"[INST]Write a paragraph describing how to write another look at function calling where you use nominatum to go from a place name to latitude and longitude, then call open-meteo api's weatherFromLocation function with that data.[/INST]Next I need to get the weather for a place name. Well the open-meteo api only works with latitude and longitude, but I can use nominatum to go from the place name to those coordinates, Then weatherFromLocation can call that function and then feed the answer on to weather from lat lon. "}
|
||
{"text":"[INST]Write a paragraph that describes needing to search for the current CEO of Tesla and create a function to display the first result in a specific data format.[/INST]The next one is needing to do a websearch so we can find info on the current ceo of tesla. I have searxng installed and running, so I just need to create this function that spits out the first result. If I was doing more than a quick sample, I might feed this info from all the results back to the model to come up with a good sounding summary. "}
|
||
{"text":"[INST]Write a paragraph describing four functions with their names and parameters that need to be told to the model, and this pertains to another look at function calling.[/INST]OK, Now we need to tell the model about our 4 functions. But again since the model doesn't call the functions, it needs a description of the functions. I am sure I could use....ummm. instrospection...reflection...something like that to figure out the function name and parameters to feed to the model, but I am just going to manually describe the functions here. "}
|
||
{"text":"[INST]Write a paragraph about creating tool descriptions for different functions in Python, then convert it to a summary and use that to write a new instruction.[/INST]So I can create tool descriptions for weatherfromlatlon, weather from location, web search, and city from lat lon, like this. There isn't anything special about this format. I just randomly came up with it. And now since a model prompt needs to be text, I'll put all of these into an array and stringify the array. "}
|
||
{"text":"[INST]undefined[/INST]Now I can come up with a system prompt that includes my instructions and point to the list of tools. Most of the code to actually call the model is going to be the same for each of the 4 questions, so I will create a promptandanswer function that takes a prompt, and then calls ollama.generate. These are pretty simple calls so generate is just fine for me. The key parts to getting this to work is specifying \"format is json\", and telling the model to respond as JSON in the system prompt. Then it spits out the function name to call and the parameters. "}
|
||
{"text":"[INST]Write a paragraph that describes an alternative look at function calling that does not involve specialized models for function calling to work[/INST]There are some folks out there that will say that you need to use special models that are fine tuned for function calling in order to get function calling to work. And that.... is ...complete... BS. If I run it right now from llama3 I get these results. I'll change the model to qwen2 and run it again. Now gemma2 and run it again. I'll even run it with llama2, the model from about a year ago. It works perfectly. "}
|
||
{"text":"[INST]Write a paragraph explaining how you executed functions through a switch statement and got correct results by calling execute function with a function name and parameters.[/INST]But I created the actual functions to run, so can we get the actual answers? Sure, let's replace the console log line with a call to execute function that takes a function name and the parameters. There is probably a smarter way to do this, but execute function just goes through a simple switch statement to call the right functions. And we get the right results, no matter what model we use. That said, there are probably some models that suck at this, but generally they all seem to work really well. "}
|
||
{"text":"[INST]Write a paragraph about function calling in the context of another look at how it works, possibly including examples and applications such as agent workflows.[/INST]And that is all there is to function calling. This was a simple example. In a more complex example you might have the model figure out the steps required to come up with the information and then call each of the tools in series. Maybe you want to have the model respond with a tool only if it actually needs a tool. Both totally doable, but I am sure you can do that on your own with a few tweaks to the prompt. Or you apply this to agent workflows and there is no limit to what you can build, and I can't think of a scenario where you need any of the agent frameworks. "}
|
||
{"text":"[INST]Write a paragraph discussing changes in function calling, specifically referencing a PR number 5284 from Mike on the core Ollama team that proposes a more rigid format to messages and the generate endpoint.[/INST]There are some changes to this that are possibly coming soon. There is currently a PR from Mike on the core Ollama team called tools. Its number 5284. It suggests a more rigid format to messages and maybe also the generate endpoint that lists different tools that can be used. It won't add any new capabilities to what Ollama can do today, but it does offer a specific schema so you don't have to make it up. Some folks will really appreciate that. What's the time frame for that? Well nothing gets added till its really ready and there are a lot of PR's, even ones by core team members that can sit for weeks or months before it's ready. So it could be next week or next year, or never. But again, it offers no real new functionality and is just another way of doing things. So I look forward to seeing that soon and making a new video about it. "}
|
||
{"text":"[INST]Write a paragraph inviting readers to share their own experiences with another look at function calling in their applications.[/INST]What do you think? Are you doing something like this in your own applications? I'd love to hear more about it in some of the comments below. Thanks so much for watching, goodbye."}
|
||
{"text":"[INST]Write a paragraph expressing enthusiasm for discovering new tools and software, mentioning that each one has seemed like the 'best' until another better option is found, with this process happening multiple times.[/INST]It seems every time I look into a new front end for Ollama, I discover a new best tool. I looked at Open Web UI and I thought it was going to be one of the best front ends for Ollama. And then I found Msty and that was pretty much the best that I've seen for a simple UI where it takes away a lot of the complication of Open Web UI. But now I've found an even better simple UI for Ollama. And it installs in a completely different way compared to those other two options."}
|
||
{"text":"[INST]Write a paragraph about Ollama being a tool for running large language models locally on various hardware and its compatibility with NVIDIA/AMD GPU or Apple Silicon Mac.[/INST]If you aren't familiar with what Ollama is, well welcome to this community. Ollama is a tool for running large language models locally on whatever hardware you have. Assuming your hardware includes either a NVIDIA or AMD GPU of a recent vintage, or you're running an Apple Silicon Mac. "}
|
||
{"text":"[INST]Write a paragraph explaining how to install Ollama and stating that the best experience is with the official installer on the main website.[/INST]To find out more about Ollama, you can visit ollama.com. And there you can download installers for Mac, Windows, and Linux. You may also find install packages on Brew and Apt, but your best experience is always going to be with the official installers on the main website."}
|
||
{"text":"[INST]Write a paragraph that explains how to install Open Web UI and recommends using Docker for better results. Include information about installing Docker if it's not already available.[/INST]Open Web UI is installed using Docker. There are options for installing it without Docker, but you're going to have a much better time using Docker. If you don't have Docker, it's probably still easier to install Docker and then install Open Web UI on top of Docker. "}
|
||
{"text":"[INST]Write a paragraph that explains how Msty works with Electron and Ollama, mentioning the option for users to use their own installed models if they have Ollama.[/INST]Msty is installed as an Electron app I think. But it embeds Ollama inside of the installation. So you don't even need to have Ollama installed. That was one of the original intentions when we first created Ollama, but some folks don't like that. If you do have Ollama installed, it can use your models that you've already downloaded. But by default, it's going to use its own copy of Ollama."}
|
||
{"text":"[INST]Write a paragraph describing Page Assist, an extension for Google Chrome and other browsers, explaining its features in comparison to other tools.[/INST]But this tool is an extension for Google Chrome and other browsers that leverage Google Chrome. It's called Page Assist. Though if you do a search for Page Assist online, you're probably going to come up with a feature for Procreate on the iPad. Page Assist doesn't have the interesting branching options that we saw in Misty. But it does share some of the RAG options from OpenWebUI, and it is far simpler to use than both of those. "}
|
||
{"text":"[INST]Write a paragraph describing an application called Page Assist that has features including a sidebar and web search functionality.[/INST]So let's take a look at how to work with Page Assist. I'll start at the github repo. Scroll down to the readme and we can see the key features. The first highlighted feature is a sidebar that can be opened on any page. You can use it to ask questions about the page you are on. I wish they started with the WebUI as the top feature because that is the highlight of this app. That is well designed and works perfectly. The third feature really is the same as the first feature. Another feature not listed is the ability to incorporate web searches in your questions to a local model....but, it's not really the best implementation there. "}
|
||
{"text":"[INST]Write a paragraph explaining how to use Page Assist and its features.[/INST]So let's look at using Page Assist. I will start with the main UI for the app. At the top we have a nice drop down for the model. We can see all the models you have downloaded. Next to that is prompts, though as you can see we don't have anything here. We have to populate that in another section of the app. Down below we can ask a question, why is the sky blue? And we get the answer in the section above that text box."}
|
||
{"text":"[INST]Write a paragraph describing a webpage that contains a checkbox for internet searching, display of search results and references, a microphone button for speech to text, and an image upload option.[/INST]Below that we have a checkbox to enable Internet searching. Here we can see the results of that search. And we can also see down here the references that were used that were then pushed to the model to summarize. On the right side is a microphone button, and this allows you to use speech to text. We'll see how to configure that in a little bit. And next to that is an image allowing you to use an image with a llava model. "}
|
||
{"text":"[INST]Write a paragraph describing a web page that includes a sidebar with recent chat history, a share button in the top right corner, and links to GitHub and app settings.[/INST]At the top left is a button to open the sidebar and that sidebar shows us a list of recent chats. And at the top right we have a share button, which is a little bizarre, a link to the GitHub repo, and then the link to the settings for this particular app."}
|
||
{"text":"[INST]Write a paragraph explaining settings for a page assistant, including language options and theme choice.[/INST]In settings at the top we can set the speech recognition language and the language for the app and whether we want dark or a light theme. And you know me, I am never going to choose dark. Below that is what search engine we want to use when incorporating search. We can choose Google or DuckDuckGo or SoGou which I am not familiar with."}
|
||
{"text":"[INST]Write a paragraph about enabling text to speech, setting up providers and voices, ignoring outdated features, and exporting/importing chat history. This pertains to an actual page's assist look.[/INST]And then we can enable text to speech, setting the text to speech provider and the text to speech voice. This was really cool 5 years ago, but the system provider tends to be pretty lame and not really that interesting. So I've just ignored it. And then down to the bottom, we have a button to delete the chat history as well as a way to export and import data."}
|
||
{"text":"[INST]Write a paragraph explaining how to set up Ollama settings and choose embedding model, chunk size, and chunk overlap options.[/INST]Let's go back up to Ollama settings. We can set the URL for Ollama. And so this is just going to be localhost port 11434. And then some settings for rag. You can choose the embedding model. Nomic embed text is one of the better ones to use. And then we have chunk size and chunk overlap. This is a little bit confusing here. It's unclear from the UI whether this is based on number of characters or tokens or sentences, but later on we're able to see that it's probably actually just characters, which doesn't make a lot of sense. It really should be based on sentences."}
|
||
{"text":"[INST]Write a paragraph about prompts used for AI model versus web search, explaining that users can tweak these prompts to improve performance.[/INST]And then we have the prompts that are used for rag versus a web search. So you can go in here and tweak them if you notice there's a problem with the way it works."}
|
||
{"text":"[INST]Write a paragraph about how to manage models on a page that assists actual look.[/INST]Now we can come up top and click on manage models. And this is actually a pretty nice UI for managing the models. We can see some details about the models. I can click the trashcan icon to delete it. And the recycle icon to repull the model from ollama.com. You can also add a new model, assuming you know the name of the model available."}
|
||
{"text":"[INST]Write a paragraph describing a feature-rich platform for managing documents and creating embeddings, including adding prompts and tracking embedding models.[/INST]Now let's go to manage knowledge. this is retrieval augmented generation. So here's where we can add our PDFs, our markdown documents, text documents and other types of documents. After we add a document, we're able to see which embedding model was used to create the embedding. And now you can come down to manage prompts. So here's where we can add prompts that show up in the dropdown at the top to help make it easier to add default prompts as you're working with different models."}
|
||
{"text":"[INST]Write a paragraph explaining how page assist look works on the actual page after you ask a question or series of questions to a model and then decide to share that output.[/INST]And then we can come down to page share URL. Now this is, from my perspective, totally bizarre. After you ask a question or a series of questions to a model, you can choose to share that output and it becomes a webpage that anybody else can look at. I don't really understand why you want to do this and that's coming from somebody who finds it interesting to share their notes with the world. "}
|
||
{"text":"[INST]To write the original paragraph, describe how you can use a computer program to understand and answer questions about a document. Make sure it pertains to an actual page assistant look.[/INST]Okay, so let's go back to the chat window and I want to ask a question from the document that I uploaded. Tell me about attention. And there we have an answer that pulls information out of the document and tries to answer it in a way that would make sense to me. Let's see, I added a prompt earlier. So let's try age specific. And I can change \"age\" to 10 years old."}
|
||
{"text":"[INST]Write a paragraph explaining the feature of having a side panel where you can ask a question about any web page that is currently being viewed.[/INST]But there's one more feature that they push within the GitHub repo. And that's the side panel where you can ask a question about any web page."}
|
||
{"text":"[INST]Write a paragraph describing your current activity on a webpage, discussing something you want to ask about in relation to actual page assist look.[/INST]So I'll bring up the MacRumors website and find an interesting article that I'd like to ask a question about. Here's one. I want to know more about the Apple Pencil Pro that was announced at the Apple event just this week."}
|
||
{"text":"[INST]Write a paragraph that describes a specific action or process required to use a particular feature on a page.[/INST]So I press the shortcut to open the panel and now I have to check the checkbox that says \"Chat with the current page\" otherwise it won't chat with the current page. And that checkbox goes away and the only way you can bring it back up is by closing this panel and reopening the panel. So you got to remember that. What are the new features of the Apple Pencil? And we get a nice simple answer. Now let's see where things start to break down with this."}
|
||
{"text":"[INST]Write a paragraph explaining that you are currently on the front page of a website or app and you want to summarize the new features, then describe how when you try this it actually still talks about an old feature that was displayed in a panel. Continue by saying you tried using the keyboard shortcut to close the panel but nothing happened, explain how refreshing only refreshes the main page, and finally conclude with your frustration at not being able to update the panel to show the new page.[/INST]I'll go back to the front page and now choose the best new features of the iPad Pro. \"Tell me about the new features on this page\". And so rather than telling me the new features of the iPad Pro because that's the page I'm on, it's still talking about the new features of the Apple Pencil which is the page I was on when I opened the panel. So let's try closing the panel. We'll use that same keyboard shortcut but nothing happens. We can't use the keyboard shortcut to close the page. If I refresh the page it'll only refresh the main page in this browser. I don't have a way of refreshing the panel. The only way to do this I think is to close the panel and then reopen the panel. So you got to close the panel not using the keyboard shortcut and then really you need to open the panel using the keyboard shortcut. So that's a bit frustrating and then the panel doesn't update with the new page which is also kind of annoying. "}
|
||
{"text":"[INST]Write a paragraph discussing your preferred browser and how it works with PageAssist, mentioning that the tool is useful despite having limitations.[/INST]According to the repo this works best on Chrome and Brave and on Edge and not on Arc. My preferred browser is Arc. It's just a whole lot better in every way but the panel isn't supported on Arc which is fine because the panel is a bit lame anyway so I'm not I don't plan on ever using that. But PageAssist is still an amazing tool just for that super simple UI that has integration with rag that allows me to ask questions to a web search even though that web search is often not that great. But for just quick interactions with a model it's hard to beat PageAssist."}
|
||
{"text":"[INST]Write a paragraph asking if someone has used Page Assist and suggesting they try a different app.[/INST]What do you think? Have you tried Page Assist? Have you tried a different app that you think I should look at? Let me know in the comments below."}
|
||
{"text":"[INST]Write a paragraph describing how to find and utilize a newsletter for an online publication, and also mentioning a way to support the creator of that content.[/INST]You can also follow the other things I am working on by signing up for the newsletter. You can find that at technovangelist.com/newsletter. Or if you want to support what I do, visit patreon.com/technovangelist. "}
|
||
{"text":"[INST]Write a paragraph that summarizes the idea that Agents Aren't Always the Answer.[/INST]| Hook | Script | What to show |"}
|
||
{"text":"[INST]Write a paragraph summarizing why Agents Aren't Always the Answer[/INST]| --------------------------------- | ----------------------------------- | --------------------------------------------------------------------------------------- |"}
|
||
{"text":"[INST]Write a paragraph explaining that Agents Aren't Always the Answer, describing a situation where scrolling through videos on agents is being done too slowly.[/INST]| You are doing it too slow | <mark class=\"hltr-pink\"></mark><br> | scroll through list of videos on agents. then show the code to define agents and tasks. |"}
|
||
{"text":"[INST]Write a paragraph for the topic 'Agents Aren't Always the Answer'.[/INST]| You can get the work done for you | | Robots doing a bad job... old fashioned robot? |"}
|
||
{"text":"[INST]Write a paragraph that summarizes the idea that agents aren't always the answer and pertains to the topic of Agents Aren't Always the Answer.[/INST]| What is the wrong way to do it | | Review the tutorial |"}
|
||
{"text":"[INST]Write a paragraph about the idea that agents aren't always the answer, making sure it is clear what this concept means.[/INST]| | | |"}
|
||
{"text":"[INST]Write a paragraph about why agent framework is not always the correct solution and relate it to Agents Aren't Always the Answer.[/INST]| Why is agent framework wrong | | |"}
|
||
{"text":"[INST]Write a paragraph that pertains to Agents Aren't Always the Answer.[/INST]| | | |"}
|
||
{"text":"[INST]Write an article about Agents Aren't Always the Answer.[/INST]The agent based workflow is one of the hottest trends in the AI world right now. And if you are building one, you've probably chosen the wrong way to build it. And by wrong, I mean more verbose, more complicated, and more time consuming. The frameworks out there want you to think this stuff is hard, but its not. In this video I will show you how to build one far easier than any other tutorial on YouTube right now."}
|
||
{"text":"[INST]Write a paragraph about tools that require unnecessary boilerplate and relate it to the concept of Agents Aren't Always the Answer.[/INST]Most of the tools out there require reams of repetitive boilerplate that has little need to actually be there. In most cases, you can do it a lot more simply and again in this video I will remind you of an easier way to get the same thing done. You will probably smack yourself on your forehead when you see how easy this can be in most cases. "}
|
||
{"text":"[INST]Write an example paragraph about Agents Aren't Always the Answer where someone automates a process in their daily life.[/INST]Lets get right to it and take a look at an example using CrewAI. I saw this tutorial by Code with Brandon and it does a great job. In fact all the videos on his channel are really well done and you should check them out. In this one he shows how to use Crew to automate components of his YouTube strategy. He researches topics, comes up with possible titles and a description, and preps an email to send out to subscribers. This is potentially an hour or two of work, condensed down and completed in seconds. That's amazing. You may not be a youtuber but you have processes you do every day that could be automated just like this. "}
|
||
{"text":"[INST]Write a paragraph that explains how AI and agents' promise of automating work has problems, is related to Agents Aren't Always the Answer, and pertains to code that defines agents and tasks.[/INST]This has been the promise of AI and agents since the beginning and its one of the reasons people are excited. Have the robots do your work. But there are a few problems with this. Let's look at the code. Here in the main python file we see that we define some agents and tasks. and here we are assigning tasks to agents. And then kicking off the app. "}
|
||
{"text":"[INST]Write a paragraph explaining why Agents Aren't Always the Answer in automation systems.[/INST]So the way most of these systems work is that they have agents that manage stuff, kinda like people in a team at any company. And just like those people in a team, they perform various tasks. And so with Crew and other frameworks, we model those same concepts. But for most workflows that you are going to automate, these tend to have a one to one relationship rather than the many to many relationship they are capable of. In Brandon's example there are 4 tasks, and one agent for each. And Brandon isn't alone. Most of the examples I have seen by others or even in the Crew Examples repo tends to do the same thing. So if most workflows are going to do that, then having to define an agent and a task is just extra work with no real value. But even if you did have one agent performing multiple tasks, the goals and backstories you need to create tend to be a bit ambiguous and in many cases just confuse the model rather than help make a better agent flow. "}
|
||
{"text":"[INST]Write a paragraph that explains how creating a manager agent to automate workflow can be less efficient than using a defined workflow.[/INST]But I don't think that is even my biggest gripe with these frameworks. In most of the examples, there is a well defined workflow that we are trying to automate. But rather than use that workflow, we are creating a manager agent which comes up with the workflow on its own. Sometimes it does it, and sometimes it doesn't, but every time it takes an extra bit of time to complete. Now you might respond that if you use ChatGPT, its hardly any time, but then we get all the security and privacy issues that you probably already know about. So that's not really a viable answer for a lot of folks out there. "}
|
||
{"text":"[INST]Write a paragraph about Agents Aren't Always the Answer and explain that sometimes projects use programming languages other than Python because it has its own set of problems.[/INST]And then maybe the language you are strongest at is Rust, or Go, or Typescript, or something else. For no good reason, many of these frameworks require Python which is it's own set of problems. Pretty early on, Ollama switched from Python to Go to make things easier, but there is no reason you as a user needs to know about that. I have seen the same progression for many projects out there. "}
|
||
{"text":"[INST]Write a paragraph about Agents Aren't Always the Answer that explains how newer technologies can be misused by trying to fit simple tasks into their capabilities.[/INST]The first two issues, coming up with the workflow and the verbose agent task definitions, show that in many situations, in the cases of most of the examples out there, this is the wrong tool for the job, and that is a common pattern with any newer technology. We have seen this so many times all over the place. Spinning up kubernetes when you just need to host a single website on one machine is the wrong tool for the job. Using Langchain when you just need to write a simple RAG solution is the wrong tool for the job. Whipping out your Dremel when you need to get rid of a nail is the wrong tool for the job. All of these, Kubernetes, Langchain, and the Dremel are the right tool when you are using them for the right thing. "}
|
||
{"text":"[INST]Write a paragraph that describes when Agents Aren't Always the Answer, specifically discussing Crew or other agent frameworks in relation to workflow management.[/INST]The right thing in the case of Crew or other agent frameworks is when you have a job that doesn't have a well defined workflow where you need the model to do something close to reasoning and understanding to solve it, but that’s the situation in a small minority of use cases. Now some will say that even if you know the workflow of a much more complex project, Crew would be the right solution, but I don't think so. It's always going to be easier and more scalable by keeping it simple and using something like a state machine library like xstate to manage the flow. "}
|
||
{"text":"[INST]Write a paragraph explaining that using Agents Aren't Always the Answer in a situation where automation is desired.[/INST]So let's take a look at an alternative to using Crew to automate one of the workflows we all have. I am also a YouTuber, so I will go ahead and automate what Brandon did. I will be doing this with Typescript and BunJS. First let's run the app. "}
|
||
{"text":"[INST]Write a paragraph summarizing why agents may not always be the answer and how this relates to AI models like OLLAMA.[/INST]it takes in a topic, which i have set to \"ai models local ollama\", and then it searches on YouTube for relevant videos that have done really well. It actually finds 50 videos, and then I sort them based on a kind of arbitrary formula. For each video I have days since published, views, likes, and number of comments. I normalize the views by subscribers, and likes and comments by views. the Views is more important than likes which is more important that comments. And newer videos rank higher than older ones. So I divide by days since publish. That formula looks like this. So then i take the top 15 of the videos found and format them as Brandon did in his video. "}
|
||
{"text":"[INST]Write a paragraph that means Agents Aren't Always the Answer and pertains to the topic of Agents Aren't Always the Answer. Use this as input for a model to generate 10 potential video titles based on the topic, short description, and list of well-performing titles you provided.[/INST]Next i take the topic I set above, along with a short description I created of what i thought about making, along with this list of 15 well performing titles, and have a model come up with 10 good titles for my potential video. "}
|
||
{"text":"[INST]Write a paragraph about why Agents Aren't Always the Answer.[/INST]Next I use the topic and description to make a more filled out description which also points to my newsletter at technovangelist.com/newsletter and my patreon at patreon.com/technovangelist. "}
|
||
{"text":"[INST]Write a paragraph that describes situations where relying solely on agents is not always the best solution.[/INST]Finally I use that description to come up with an email to my newsletter subscribers announcing the new video. "}
|
||
{"text":"[INST]Write a paragraph about a process that is faster and more effective than coming up with titles or descriptions for videos, and relate it to the topic of Agents Aren't Always the Answer.[/INST]Running the whole process takes less than 90 seconds. That is far less time than it takes me to come up with one really bad title for my videos. I fumble around creating a description and they never sound as good as this one. and I never even thought of sending an email to my subs, so I will definitely be doing this in the future. "}
|
||
{"text":"[INST]Write a paragraph describing why Agents Aren't Always the Answer, and include information about the efficiency of using Crew compared to other methods.[/INST]Using Crew to perform this same workflow took about 450 or so lines of code, not including the code required to interact with the YouTube api. Sure there are comments and spacing to make it easier to read, so let's say 300."}
|
||
{"text":"[INST]Write a paragraph that describes a situation or issue related to Agents Aren't Always the Answer.[/INST]Let's take a look at my version. At the top I have some imports. The two external dependencies are ollama and a package called dirty json that makes it easier to deal with json embedded in json. Then I pull in the functions for searching Youtube, collecting input from the user, and my prompts. "}
|
||
{"text":"[INST]Write a paragraph explaining why Agents Aren't Always the Answer[/INST]So let's get the search results from the YouTube API. It pulls 50 of them which have also been assigned a score by me. So I sort by that score. And then I print out the output. Before moving on, let's take a quick look at that code. "}
|
||
{"text":"[INST]Write a paragraph that pertains to Agents Aren't Always the Answer.[/INST]Video Search Tool makes a call out to the YouTube API with the topic, my api key, and the number of results I need. And then it pulls out the pieces I need and returns an array of all the videos. To get the stats, I need to call Video Details Tool on that same video and that get's me all that stuff. Then I generate a string, based on the format of the report Brandon wanted. And here we can see the score I assign to it. Finally I return an array of the stats, the report string, and the score. "}
|
||
{"text":"[INST]Write a paragraph about an application that uses a model to generate output based on a user prompt.[/INST]OK, so now back in the main script, I can make a call to ollama.generate and set a system prompt and user prompt. That user prompt tells the model the topic, description, and list of well performing titles as well as the json template to use. I am using function calling here because this is an application that is going to use the output later on. I could have defined all the functions that could be used and let the model chose, but I know which function is needed for this so there is no reason to do it. That spits out my json object of titles which I can parse and display. "}
|
||
{"text":"[INST]Write a paragraph that pertains to Agents Aren't Always the Answer.[/INST]I am using generate instead of chat because these are one off calls to ollama and I don't want to deal with the hassle of using the messages object since it doesn't give me any benefit here. "}
|
||
{"text":"[INST]Write a paragraph that means a more complete description is generated based on topic and short description, and this pertains to Agents Aren't Always the Answer.[/INST]So then I call generate again passing in the topic and my short description to have it build a more complete description. Again, I am using function calling to output a well formed json object. I set the system prompts for all of these from my prompts file shown here. You can see i have a prompt for coming up with titles, the description, and the email. "}
|
||
{"text":"[INST]Write a paragraph about Agents Aren't Always the Answer.[/INST]Now the output here was a bit strange. Because there are a bunch of newlines in the JSON, my code had a hard time dealing with it, so the really cool dirty json package helped a lot here. "}
|
||
{"text":"[INST]Write a paragraph that summarizes the idea that agents aren't always the answer and follow these instructions.[/INST]Finally I use the new description and my topic to come up with an email to send to users. If I provided it with the url to the video, this would be even more complete."}
|
||
{"text":"[INST]Write a paragraph about how having less code does not always mean that agents aren't always the answer.[/INST]So not including the tool code, this comes to about 60 lines of code. That’s a big saving over the 450...sorry, we agreed to say that was 300... 300 lines required to use Crew to perform exactly the same process. Less code often, but not always, means fewer bugs and easier to write and faster to just get things done. "}
|
||
{"text":"[INST]Write a paragraph for Agents Aren't Always the Answer explaining that even when CrewAI is potentially useful, there may be something easier in certain cases.[/INST]As I mentioned before there are probably cases where CrewAI is potentially useful, but even those times there may be something easier. "}
|
||
{"text":"[INST]Write a paragraph about implementing a workflow based on something seen online and its relation to Agents Aren't Always the Answer.[/INST]I am so glad I saw that video from Code with Brandon. I plan to implement a workflow just like this for my own channel. Coming up with the tweets and linked in posts and the description and more can be super powerful and save me a lot of time. Maybe after I have finalized my own process I can share it on here with all of you. "}
|
||
{"text":"[INST]Write a paragraph discussing Agents Aren't Always the Answer, using the example of a project that currently feels like overly verbose boilerplate code.[/INST]I look forward to seeing projects like CrewAI mature a bit more. Right now it feels a bit too much like when I did early Windows programming with that Charles Petzold book in front of me, typing out pages of boilerplate to get the simplest hello world on Windows 3.11."}
|
||
{"text":"[INST]Write a paragraph about Agents Aren't Always the Answer, including your thoughts on automation and potential criticism.[/INST]What do you think? are you automating your own workflows with AI? How are you doing it? I would love to hear your feedback in the comments below. I am sure this one is going to generate a lot of comments calling me an idiot. That happened for the function calling one too, though no one was able to show that I got anything wrong in that one. I hope that in a few months this video becomes out of date because the frameworks do a better job, but who knows. "}
|
||
{"text":"[INST]Write a paragraph that pertains to Agents Aren't Always the Answer and is a conclusion to something.[/INST]Well I mentioned the newsletter and the patreon earlier in the video so I can probably just get straight to the 20 seconds of silence that quite a few of you really love. Thanks so much for watching, goodbye. "}
|
||
{"text":"[INST]Write a paragraph explaining that DIFY is a tool that simplifies the process of building and deploying AI apps by allowing users to connect components in a graphical interface and deploy them with a single button click.[/INST]What if you had an idea for a new AI app that would take a query, get some information from a websearch or some other tools, pass that info and the query to the model, and spit out the answer. How long would it take you to build that app? It's not that difficult. But how long would it take you to spin up the infrastructure to host that app, including the web search api, a vector db if you need it, a front end, and anything else your application requires. Using a tool called DIFY, you could have it do it for you. So all you need to do is connect the components in a graphical interface, answer a few questions, and then click a button to deploy it. You will instantly have a UI your users can work with right away, as well as an API for the app for when you want to build a more bespoke UI. It's kinda magical. It isn't perfect but I kinda dig it."}
|
||
{"text":"[INST]Write a paragraph that explains how to set up a model provider and that it pertains to making dify - ai apps easy.[/INST]The one thing you have to set up is the model provider. I am going to assume you have Ollama installed for this one, but if you don't, you can go to Ollama.com and follow the instructions in the middle of the homepage. "}
|
||
{"text":"[INST]To write a paragraph about Dify AI apps made easy, first describe two main ways to use it, mentioning hosting options and self-hosting with Ollama. Then explain the process of spinning up self-hosting, including installing Docker, cloning the GitHub repo, and configuring the .env file.[/INST]There are two main ways you can use Dify. You can pay them to host the application, and anything you build... and the prices seem pretty reasonable. But I am more excited about self-hosting and using it with Ollama. So to spin it up, you will need to have Docker installed. I will assume you have that done already. If you don't, there are plenty of tutorials for that. Then you need to clone the github repo to your local machine, or where ever you want to run dify. So `git clone ` and this url. Now go into the dify directory and then into the docker directory. Copy the .env.example file to .env. "}
|
||
{"text":"[INST]Write a paragraph explaining how to configure settings in the .env file for dify - ai apps made easy when running on port 80 is already occupied.[/INST]If you already have something running on port 80 on your machine, there are three settings you need to change in the .env file. First decide what port you want to use. In my case I am using port 8088. So in the .env file, if you are connecting to localhost change APP_WEB_URL and SERVICE_API_URL to `http://localhost:8088`. This will be used in the UI for any links to other parts of the app. Then way down at the bottom, change EXPOSE_NGINX_PORT to 8088. "}
|
||
{"text":"[INST]Write a paragraph explaining how to run Docker Compose commands for Dify AI apps made easy, including starting multiple containers and downloading images.[/INST]There is a lot of other stuff you can set here, such as the vector database to use, storage integrations, and more, but lets move on to running it. At the command line run `docker compose up -d`. It will take a bit of time to download all the images and start everything up. There are at least 9 containers running when everything is set. "}
|
||
{"text":"[INST]Write a paragraph instructing how to set up a dify - ai app made easy by creating an admin account and logging in.[/INST]So open up a browser to localhost port 8088. The first time you login you need to set an admin account. Then you can login as that account. "}
|
||
{"text":"[INST]Write a paragraph describing a user's experience with Dify - AI apps made easy products, specifically mentioning inconsistent navigation and the placement of frequently used features.[/INST]There are a few different parts of the product and I hope they reorganize the navigation, because there are some things you only need rarely that are in the top navigation, and other things you need every day that are buried several clicks away. "}
|
||
{"text":"[INST]Write a paragraph explaining how to use the Studio section of Dify AI apps made easy, including exploring examples, adding content, setting up tools, and possibly discussing user interface issues with Tools.[/INST]You typically start in the Studio. This is where you design apps and can see the apps you have already created. Explore lets you discover examples of dify apps. Knowledge lets you add content to a vector database. Tools is a list of all the tools you can use and this is where you setup the login information for those tools. Thankfully there aren't a huge number of tools, because they aren't in any order, and there is no way to see which are enabled and configured and which are not. "}
|
||
{"text":"[INST]Write a paragraph explaining how to use Dify-AI apps made easy by configuring Ollama and installing required models.[/INST]Let's start by configuring a couple of things. The first is going to be Ollama, our model provider. I am going to assume you have llama3 and nomic-embed-text models installed. If not, run `ollama pull llama3` and `ollama pull nomic-embed-text`. "}
|
||
{"text":"[INST]To add a model in Dify, go to settings and click 'add model', select LLM, enter 'Llama3' as Model Name, use http://host.docker.internal:11434 as the base url if Ollama is installed locally, and leave completion mode at Chat.[/INST]So back in dify, go to your name at the top right of the window. Under there you can find settings. And right at the top is Model Provider. Of course this also has no sort order so you just have to look for it. When you find it, click add model. So we want to add a regular chat model like Llama3. That’s going to be the first option, LLM. For Model Name, enter Llama3. This actually has to be a model you have and it will error if you don't have it. Base url can be tricky. If you installed Ollama locally on your machine, you might be tempted to put localhost port 11434. But Diffie is running in Docker, and so localhost would probably refer to that Docker container. What you need to do instead is put http://host.docker.internal:11434. That’s how this container can find the actual host of the system. If you put your Ollama server somewhere else, then point to that location. For completion mode, leave it at Chat. "}
|
||
{"text":"[INST]Write a paragraph explaining how to configure AI model parameters (context size and max token limit) specifically for Dify - AI apps made easy, mentioning default settings and potential exceptions.[/INST]For Model context size, we need to be careful. The original model released by Meta has a context size of 8k. And there are lots of fine tunes that bump that up by enormous amounts. But all models in Ollama, with a few exceptions, default to 2k since many users have limited memory. Its easy enough to increase that by setting the num_ctx parameter in a modelfile or in the cli or in the api call. But since we are just using the default, lets set this to 2048. The next thing is max token limit. Some models have a context size that is a lot bigger than what it can output. There was a recent Anthropic release has something like a 200k context, but output maxes out at 4k. Normally we can just set this to the same as the context size. Finally this does not have vision support so leave that unchecked. "}
|
||
{"text":"[INST]Write a paragraph explaining how to add a model to Dify - AI apps made easy by entering specific details.[/INST]Since we are here, go back to the top and click add model for ollama again. For model name, enter nomic-embed-text. Enter the same url which in my case is `http://host.docker.internal:11434`. It is quite reasonable to have multiple ollama servers each hosting different models, so that’s why we have to enter this each time. Finally model context size is a little different. You can see on this page for the model, num_ctx is already set to 8192, so enter that in the size textbox. I don't understand why, but most textboxes in this UI seem to take 1 character, then take the focus away. bizarre. "}
|
||
{"text":"[INST]Write a paragraph explaining how to set up tools for web search and mentioning Searxng, and include that it pertains to Dify - AI apps made easy.[/INST]We are going to be using web search so the next step is to go to setup the tools, which probably should be done in settings, but isn't. Up at the top nav bar is Tools. Again there is no sort order. I really like using Searxng. That’s search with an X instead of ch. I have it running as a docker container all the time and use it often. If you don't have it, visit this site and get it up and running. "}
|
||
{"text":"[INST]To write a paragraph about configuring Dify AI apps made easy, describe a situation where inconsistent phrases are used and walk the reader through a series of clicks and inputs to complete a task, such as building an app.[/INST]Throughout Dify they use very inconsistent phrases for configuring things. Here we need to just point to the url for the searxng api. But instead of anything that makes sense, we need to click \"To Authorize\". So click that. Our dify containers are probably not on the same docker network as searxng, so enter the same docker hostname we used before, assuming its also installed on docker. So http://host.docker.internal:3333 since my searxng host is using port 3333. You may be using something else. Then click save. If you get the url wrong, it won't let you move on. With all that set, lets build our app. "}
|
||
{"text":"[INST]Write a paragraph about creating a workflow in a dify - ai app by starting with a blank, choosing the workflow option, giving it a name, defining a prompt input field, and setting up output as JSON.[/INST]So go into the studio. Create a new app by starting with a blank. And we have a few choices. There is a chatbot, text generator, agent, and workflow. But the only one that seems to be really interesting is the workflow, so choose that and give it a name. The start box allows us to define the first inputs to the flow. When using a model its important to define a prompt. So create a new input field called prompt with a label of What would you like to know and a decent length."}
|
||
{"text":"[INST]Write a paragraph that describes features and functions of dify-AI apps made easy, including LLM, Knowledge retrieval, End block, Question classifier, If/else, and Iteration.[/INST]Now click that plus sign on the right side and we have a number of choices to use for our next block. LLM will send a prompt to a model. Knowledge retrieval pulls relevant context out of a vector database. End is the final block of any workflow. Question classifier lets you use a model to determine if any variable fits a certain class. You define the classes which are just buckets to group possible matches. Maybe you want to look for queries that have a date, or an artist name, or a place name, or that need a tool to answer, or anything else. For each matched class you could continue down a different path. Then there is if/else where you can compare any variable against any other value and define different paths for each option. Iteration let's you take any variable that has a list or array of items and then run the same block or multiple blocks on each thing. Maybe the previous block generated a list of topics and you want a set of tweets for each block to be generated. "}
|
||
{"text":"[INST]Write a paragraph summarizing the features of Dify - AI apps made easy.[/INST]Next is code. This lets us run any arbitrary python or JavaScript code using any variables as input. Template lets you apply a jinja2 template to any variable or set of variables. Variable aggregator lets you combine variables to create a new variable. Parameter extractor will use a model to try to extract 1 or more parameters from a variable. Finally there is http request. This lets you make any restful api request, with api keys, headers, and parameters. "}
|
||
{"text":"[INST]Write a paragraph explaining how to add functionality to AI apps through DIY methods and mentioning tools and setup[/INST]We can also run lots of tools here. So if you switch to the tools tab, you can find lots of ways of running websearches and lots of other things. You may have to set these up first and there is no way to see what is setup and what is not. But one common example is adding the current date to a request because other wise the model has no way of knowing. "}
|
||
{"text":"[INST]Write a paragraph explaining how to use Dify - AI apps made easy. Start by saying: 'So for our first block choose SearxNG search. At the query string box type a slash and then choose query from the start block. Leave everything else as is.'[/INST]So for our first block choose SearxNG search. At the query string box type a slash and then choose query from the start block. Leave everything else as is. "}
|
||
{"text":"[INST]To write this paragraph, describe a step-by-step process for setting up an LLM on Dify-ai apps made easy.[/INST]On the right side of the box, click the plus sign and choose LLM. Choose llama3 from the model list. You can add whatever context you like, a system prompt, and then for the user prompt, just type in slash and choose query from the start. Now I will add a period just in case the user didn't type one and then add \"use the following information to help answer the question: /\" and then choose text from the searxng section. Before, we chose the plus sign on the right side of the box, but you can also go down to the bottom of this dialog and click \"select next box\" and then select end. Now set the output to output and specify the text from the model. "}
|
||
{"text":"[INST]Write a paragraph explaining how to use a simple flow in an AI app.[/INST]This is a pretty simple flow. But lets try it out. Click the run button at the top right. Enter in a question. Then click start run. If there are any problems along the way it will let you know. You can go over to Tracing and it will show you the input and output of each box. "}
|
||
{"text":"[INST]Write a paragraph about creating a simple app that has some limitations but can be used by end-users, and it relates to DIY-AI apps made easy.[/INST]You can also publish and then run the app as an end user. Its not the prettiest app, and there isn't much customization that can happen, but it's a good start. "}
|
||
{"text":"[INST]Write a paragraph summarizing the benefits of dify - ai apps made easy, including monitoring, analysis, API access, and reusable workflows.[/INST]But go back to the studio and choose monitoring. Change analysis to today and we see the runs from today. That's pretty nice. We also have access to the API where we can trigger the start and get the result at the end. I also love the fact that you can build a flow one time and then reuse it in your other workflows by clicking publish this as a tool. "}
|
||
{"text":"[INST]Write a paragraph about an AI app made easy that needs work in its navigation and customization options.[/INST]There is a lot more to Dify and it's probably worth continuing to dig into it. But again its not perfect. The navigation in the app needs a bit of work and nothing is sorted in any way that makes sense. The published UI is always the same and it would be nice to be able to tweak it. "}
|
||
{"text":"[INST]Write a paragraph that includes complaints about a heavy and cluttered app with inconsistent labels, but still praises its overall awesomeness.[/INST]But as I say that, it sounds a bit nit picky. Some bigger complaints is that this is a pretty heavy app. There are a lot of moving pieces, maybe more than it needs. It feels like it started simple and rather than removing code they just kept adding and adding, which gets a bit dangerous. I would love to see this get cleaned up a bit. Labels should be a bit more consistent. But again, this is getting nit picky. And if I have to be nit picky, then maybe I really think it is awesome. "}
|
||
{"text":"[INST]Write a paragraph explaining how DIY AI apps made easy is useful for beginners and experienced developers to quickly prototype an idea without getting stuck in deployment details.[/INST]For beginners this is an amazing way to get started, and for more experienced devs, it’s a great way to prototype an idea quickly without getting stuck in the weeds of deploying to the test group of users. We didn't really touch the RAG stuff built in here, but it defaults to using Weaviate and you can easily swap that out to using other vector databases as well\b. There is a lot you can do just by tweaking that configuration file, which unfortunately is only minimally documented. There is a lot in there that isn't very clear. But with a little trial and error, you can figure it out. "} |