SUBSCRIBE NOW
IN THIS ISSUE
PIPELINE RESOURCES

Efficient Use of AI

By: Mark Cummings, Ph.D.

The initial exuberance over using GenAI in business and government is running into unexpectedly high online AI services bills. This is leading to a recognition that not all uses of AI are economically justified. This recognition is a sign of maturity. Now, the question becomes: how do you achieve AI efficiency?

Return on Tokens

The question is essentially a question of return.  Return on fees paid to service providers and the cost of staff time.  In this context, return means how much benefit you get from how much it costs. Online GenAI systems are billing based on tokens. A token is a small word (image, sound, etc.) or part of a larger word (image, sound, etc.). In online commercial systems, tokens are generally counted, and that is how the user is charged. Input to an AI is generally prompt language, data, background information, etc. Input is broken down into tokens. AI processing can be thought of as the AI talking to itself.  This conversation creates more tokens. Then, the result is output, creating the final set of tokens.

Charging for retail users is generally through a subscription, which allows users a certain limited number of tokens per unit of time. Business accounts can be done the same way, but are more likely to be charged per number of tokens used. This second way is how businesses are being surprised by their bills. Staffs are using far more tokens than anticipated. are using far more tokens than anticipated.

Thus, the question of efficiency of use of AI boils down to how many tokens plus how much staff time you are using to get what value of result. This is becoming known as Return on Tokens (RoT). There are essentially two sides of RoT: the cost side and the benefit side.  Each can be pursued independently. However, they often have to be considered together. 

The most important thing when working with AI’s is knowing exactly what you want to achieve. Knowing what you want to achieve has two aspects: prompt language and architecture.

Prompt Language

Independent of everything else, the more accurate and complete your prompt language is, the more efficient you are going to be in using AI.

If you ask a GenAI system to give you something described only generally. The AI will give you something. It will try to figure out what exactly you need. But you are not likely to get what you want. When working in a chat mode, this tends to lead to long, expensive iterations. Iterative interactions that consume more and more tokens and more and more of the user’s time. Both expensive.

In an Intelligent Agent (IA) application, prompts that are not completely accurate will cause behavior problems in the IA. If you catch it in development, it will create a string of iterations using more tokens and taking more staff time. A bigger issue is if you don’t catch the problem in development and it leads to aberrant behavior in the field.

Struggling with inaccurate prompt language can lead to multiple versions of an AI running. Sometimes, it is easy to forget to turn off one of the previous versions and a Ghost IA result. Ghost IA’s consume resources that can be expensive and can pop up in the field and cause hard-to-explain trouble.

In an organization, you are likely to have a wide variety of language capabilities. Some will be able to be naturally exact and complete. Others will be able to learn to become accurate and complete. Some just don’t have that skill. Training can, at least, make everybody sensitive to the issue.

One way to improve efficiency in this area is to use an AI to help craft prompt language.


FEATURED SPONSOR:

Latest Updates





Subscribe to our YouTube Channel