Software und Tools – ID's Blog

Juni 5, 2025Juni 5, 2025

5-Sekunden-KI-Videos mit Bing Video Creator

Dank des Heise-Artikels vom 4.6.25 zum neuen Bing Video Creator (https://www.heise.de/news/Microsoft-spendiert-Bing-einen-KI-Video-Ersteller-10425495.html) im Folgenden eigene kostenlos erzeugte Testergebnisse als LOOP-Videos auf meinem Webserver.
Die in Bing verwendete Komponente Sora ist bei OpenAI schon länger erhältlich, allerdings dort lediglich über ein kostenpflichtiges Abo (s.a. https://openai.com/de-DE/sora/).
Microsoft bietet über die Bing App nun die Erstellung von Videos im Seitenverhältnis 9:16 in 480p und zwar mit einer Dauer von 5 Sekunden Länge.

Waschbär-Motive finde ich sehr geeignet zum Testen von KI-Anwendungen, daher auch hier meine entsprechenden Prompts.
Die KI ist – wie bei Bing bzw. OpenAI üblich – gestaltend tätig, ohne dass man selbst jedes Detail angibt. Bei Videos ist daher spannend, welche Bewegungen von der KI akzentuiert wurden. Noch geht das Erstellen eines Videos nur aus der Bing-App heraus und nicht über die Desktop-Version, doch das wird sich sicherlich bald ändern. Auf die jeweiligen Ergebnisse musste ich tatsächlich jeweils mehr als 2 Stunden warten, aber das habe ich für meine kostenlosen Tests gerne in Kauf genommen.

Versuch 1: Fishing Raccoon

„a raccoon who goes fishing with a sun flower in a pond in front of the Louvre. evening light“

Die Umsetzung finde ich gelungen mit Bewegungen des Wassers und Drehbewegung des Waschbärs bei entspanntem Ambiente.
Die Prompt-Idee habe ich für diesen Sora-Test recycelt, aber ohne Hommingberger Gepardenforelle (s.a. https://www.heise.de/foto/galerie/suche/foto/?keyword=gepardenforelle).

Versuch 2: Juggling Raccoon

„a raccoon at the beach juggling balls during a thunderstorm while a herring gull watches“

Die Möwe fehlt leider im Ergebnis, aber das Video zeigt einen sehr hektischen Waschbären, was gut zum Gewitter-Setting passt. Er will schnell mit seinem Training fertig werden, bevor das Wetter noch schlechter wird und es nicht nur über dem Wasser regnet 😉

***

Laut Heise-Artikel werden die Bing-Videos bis zu 90 Tagen bei Microsoft gespeichert; hier die Links zu meinen 2 Versuchen:

Versuch 1: https://sl.bing.net/Mum1yNGsXQ
Versuch 2: https://sl.bing.net/iyHi5tjsY7o

Februar 15, 2025

Hugging Face AI Agents Course

AI Agents sind DAS Thema derzeit. Regelmäßig gute Links dazu bietet übrigens George Siemens „SAIL“-Newsletter https://buttondown.com/SAIL/archive
So bin ich auch auf den am 10.2.2025 gestarteten kostenlosen „AI Agents“-Kurs von Hugging Face gestoßen:
https://huggingface.co/agents-course
https://huggingface.co/learn/agents-course/unit0/introduction

Die Hugging Face Empfehlung für Teilnehmende ohne Python-Kenntnisse lautet, sich (evtl. nur) mit Unit1 „Introduction to Agents“ zu befassen und ein kleines Zertifikat zu erwerben. Am Ende des Unit1-Theorieteils gibt es die Möglichkeit, ein fertiges AI-Agent-Beispiel online in einem eigenen Hugging Face Space auszuprobieren und dabei lediglich in einer Python-Datei kleinere Anpassungen zu machen. Verwendet wurde dafür „Smolagents“, das AI Agent Framework von Hugging Face.

Meine Tests mit dem in Unit1 angebotenen AI-Agent-Beispiel
Wegen Überlastung des eingebundenen Qwen2.5-Coder-32B-Instruct-LLM habe ich schnell auf die im Skript aktivierbare technische Alternative zurückgreifen müssen. Dann klappte es gut mit Anfragen und Ergebnissen mit Darlegung der Agent-„Thoughts“ bis hin zur „Final Answer“.
Bzgl. eigenen Anpassungen der Python-Beispieldatei habe ich 2 Dinge getestet:
a) Aktivieren des bereits vorbereiteten Image-Generators durch Eintrag in der Zeile tools=[final_answer, image_generation_tool],
b) Eintragen eines kleinen Lottozahlen-Skripts als Tool
Wieder einmal war die Nutzung von ChatGPT hilfreich, um mir ggf. Code erklären bzw. generieren zu lassen, da meine Python-Kenntnisse sich in Grenzen halten 🙂

Für mich interessant: Der bereits im AI-Agent-Beispiel referenzierte Image-Generator greift auf „model_sdxl = „black-forest-labs/FLUX.1-schnell“ zurück und lieferte ziemlich gute Bilder.

Chatverlauf Screenshot 1 | Chatverlauf Screenshot 2 | Chatverlauf Screenshot 3

| | |

***************************

Abschließend noch eine Erklärung aus dem Kurs, was ein AI Agent ist 🙂

„To summarize, an Agent is a system that uses an AI Model (typically an LLM) as its core reasoning engine, to:

Understand natural language: Interpret and respond to human instructions in a meaningful way.

Reason and plan: Analyze information, make decisions, and devise strategies to solve problems.

Interact with its environment: Gather information, take actions, and observe the results of those actions.“

(Quelle: https://huggingface.co/learn/agents-course/unit1/what-are-agents)

April 19, 2024April 19, 2024

Bilder generieren mit ideogram.ai

Auf der Suche nach einer Alternative zu DALLE-3 stieß ich auf ideogram.ai – auch aufgrund des Hinweises, dass ideogram.ai sehr gut mit Texten in Bildern umgehen kann. Das Web-Tool hat mir sowohl von der Qualität als auch von der Geschwindigkeit her sehr gut gefallen. Wie inzwischen so ziemlich alle Tools, erfordert es ein persönliches Login (Google, Apple) – auch für einen freien Account mit max. 25 Prompts täglich. Hier die ideogram.ai-Dokumenation: https://docs.ideogram.ai/using-ideogram/generating-images und die Nutzungsbedingungen: https://ideogram.ai/tos

ideogram.ai-Testbericht

Mit jedem Prompt werden 4 Bilder erzeugt, von denen eines automatisch zum Cover-Bild gewählt wird. Dennoch kann man manuell einzelne Bilder aus verschiedenen Kollektionen anpinnen an den eigenen Account (s.u. Screenshot). Die erzeugten Bilder sind downloadbar im jpg-Format inkl. Exif-Daten. Die Option „share“ scheint im ideogram.ai-Bereich intern zu bleiben, daher hier im Blogbeitrag manuell hochgeladene Ergebnisse.

Standardmäßig ist „Magic Prompt“ angehakt zur Optimierung der eigenen Text-Eingabe, was man auf Wunsch aber deaktivieren kann. Ansonsten kann man noch zwischen verschiedenen Ausgabe-Formaten wählen und das ist es dann schon mit der Wahlfreiheit bei einem kostenlosen Account: Erzeugte Bilder sind „public“ und man kann sie nicht löschen. Der Editor steht auch nur für Bezahl-Accounts zur Verfügung.

Screenshot meiner angepinnten ideogram.ai-Ergebnisse

Programm-Oberflaeche (zum Vergrößern anklicken)

Einzelbilder zum genaueren Anschauen

Prompt = 3 raccoons, a cat and a waterfall. wildlife photography

Prompt = a blue heron standing on a desk in front of a keyboard and looking at a monitor. mysterious yellow light, black and white items. a small pink raccoon logo in the upper right corner.

Prompt = Create an aspect ratio of 9:16 depicting half of a face of a raccoon and the other half with text „RACCOONS IN KASSEL FOREVER“ with bold and grunge effect on the text, black and white illustration with pink background, typography.
Magic Prompt = A striking black and white illustration with a 9:16 aspect ratio, capturing half of a raccoon’s face looking intently towards the viewer. The other half of its face is replaced by bold, grungy text that reads „RACCOONS IN KASSEL FOREVER“ in a playful and rebellious tone. The illustration is set against a vibrant pink background, adding an unexpected pop of color. The overall design is edgy and artistic, with a strong typography element that adds to its visual impact.
-> Mein Prompt war inspiriert von einem Beispiel auf X (Ex-Twitter): https://twitter.com/ideogram_ai/status/1775512625851375812
Meine Formatangaben wurden durch die gesetzte Option Aspect Ratio = 1:1 übersteuert.

Prompt = Stelle den Vorgang der Photosynthese symbolisch dar.
Magic Prompt = A creative visual representation of photosynthesis, where a large tree stands tall in a sunlit meadow. The tree’s leaves are depicted as glowing, capturing sunlight and converting it into energy. A series of tiny, colorful insects and birds surround the tree, symbolizing the interconnectedness of the ecosystem. In the background, there is a golden sun and a gentle, blue sky. The overall ambiance of the image is serene, with a touch of magic and life.

Prompt = cute young woman sitting on a beach reading a book labeled „Raccoons forever“, photo realistic style, mysterious sun light, typography.
-> Der Text im Bild klappte auf Anhieb, der Text auf dem Buchtitel nicht ganz so gut …

Prompt = a learning space designed with areas for quiet study, online learning, hybrid learning, power generation via exercise bikes and a comfortable relaxing atmosphere. empty – no students.
Magic Prompt = A futuristic and innovative learning space, designed for a variety of educational experiences. In the center, there are several exercise bikes connected to power generators, providing an eco-friendly way to produce energy while students study. The quiet study area features modern, comfortable seating and natural lighting, while the online and hybrid learning zones are equipped with state-of-the-art technology and interactive displays. The entire space is empty, awaiting the arrival of eager students to explore and learn in this unique environment.
-> Meine Eingabe „no students“ wurde leider ignoriert und auch die entsprechende automatisch erzeugte Eingabe des Magic Prompts … Die Darstellung der Personen ist zudem extrem schlecht – im Gegensatz zu meinem Beispiel mit dem Chemiker.

Prompt = a typical north german landscape. one blue sheep. sunny day.
Magic Prompt = A serene, sunny scene of a typical north German landscape, featuring a lush green field with various types of tall grass and wildflowers. In the center, a single, unique blue sheep grazes peacefully, contrasting with the surrounding natural colors. The sky is a brilliant blue, with fluffy white clouds floating above. The overall atmosphere is tranquil and idyllic, capturing the essence of a perfect rural day.

Prompt = Historisches Wasserschloss im Abendlicht. Bleistiftzeichnung.

Prompt = Louvre pyramid with giraffes next to it. Evening light. Photo.
Magic Prompt = A stunning photograph of the Louvre Pyramid bathed in the golden glow of evening light. Two elegant giraffes gracefully pose next to the pyramid, casting elongated shadows on the polished floor. The background reveals a serene, open plaza with a few people admiring the striking scene. The combination of modern architecture and graceful wildlife creates a harmonious and enchanting atmosphere.

Prompt = raccoon with balloon., graffiti

Prompt = Laborraum. Ein Chemiker mit einem Reagenzglas auf dem Tisch. Fotorealistisch.

Prompt = raccon for president. poster. typography
Magic Prompt = A vibrant and humorous political poster featuring a raccoon wearing a stylish suit, holding a microphone with the slogan „Raccoon for President.“ The raccoon’s face is replaced with the American flag, emphasizing its patriotic message. The background is a cityscape with skyscrapers, and the typography is bold, modern, and eye-catching. The overall tone of the image is playful and satirical, poking fun at the political landscape.

Prompt = a classroom. van Gogh style.

Prompt = a raccoon, a giraffe and a cat looking at the sunset over a river. black & white coloring book style illustration.

Prompt = self portrait of a emperor penguin as an artist, painting on a canvas in a historical setting that echoes the ambience of a 17th century Europe.

Prompt = 3 raccoons, a cat and a waterfall. wildlife photography
Magic Prompt = A serene wildlife photo capturing three raccoons and a cat coexisting peacefully near a picturesque waterfall. The raccoons display a mix of curiosity and caution, while the sleek cat observes intently from a distance. The waterfall cascades in the background, casting a mist that shines in the sunlight. The lush greenery and rocks in the foreground give a sense of the animals‘ natural habitat.

Prompt = Bergpark Kassel, Tiger, Katze und Maus., 3d render
Magic Prompt = A stunning 3D rendered scene of Bergpark Kassel, a majestic mountain park featuring a gigantic tiger, a curious cat, and a playful mouse. The tiger stands proudly on a hill, surveying the expansive park, while the cat and mouse engage in a lighthearted game nearby. The setting is a lush, green environment with vibrant flowers and trees, creating a serene atmosphere., 3d render
-> Der „Bergpark Kassel“ war dem Tool offensichtlich kein Begriff … (im Gegensatz zum Begriff „Louvre“)

Prompt = Louvre pyramid with giraffes next to it. Evening light. Photo.
Magic Prompt = A stunning photograph of the Louvre Pyramid illuminated by the warm evening light, with a delicately blurred effect. Near the base of the pyramid, two majestic giraffes gracefully pose, their long necks reaching towards the sky. The background reveals a picturesque view of the Louvre Museum and a serene cityscape. The overall atmosphere of the image is peaceful and enchanting, blending architecture, wildlife, and urban elegance.

Januar 27, 2024

Stable Diffusion locally on MacBook Air via App „Draw Things“

It would be nice to have a local installation of the open source text-to-image model Stable Diffusion… that was my idea this weekend. As I didn’t find any convincing installation manuals and even wasn’t sure if my Windows PC’s hardware would be sufficient, a very simple article about „Draw Things“ came to my attention: https://www.unidigital.news/draw-things-ki-bilder-kostenlos-unterwegs-generieren/ Unfortunately my iPad and my iPhone are too old, but with my MacBook Air (M1,2020) the installation was no problem:
Downloading the app „Draw Things“ from the Apple Store and afterwards downloading the proposed model „SDXL Refiner v1.0 (8-bit)“. My first text prompt with default options, 1024×1024 etc. and waiting time of some minutes, was a disaster, my second prompt was a disaster and so it went… double things and not nearly what I wanted. O.k., I’m a little bit spoiled because of mostly using ChatGPTPlus and Bing Image Creator, but what was that thing? Totally disappointed, I deleted all the results.

Here is one disturbing example from today „young woman holding an umbrella“:

(click to enlarge)

Today, I thought maybe I should be modest and download/install an older Stable Diffusion model in my „Draw Things“ app and therefore chose „Generic (Stable Diffusion v2.1)“ which would result in a lower resolution image 512×512. For my prompt I chose something about cats (because, normally „cats always work“) and finally, I got results resembling a cat:

(click to enlarge)

The good results with the older model led me to the assumption, that the SDXL model was the problem and I googled… a YouTube video „Fix Double-Headed Glitches in Stable Diffusion with Kohya Hires Fix!“ helped me understand the problem (Link: https://youtu.be/SbgMwHDXthU?feature=shared). As I don’t have a local installation but the Apple Store App „Draw Things“, I can’t install any extensions, but I looked instead at the menu „Advanced“ in my „Draw Things“: There is indeed already a solution, that means a configuration option named „High resolution fix“ (description „[…] it avoids duplicate objects when generating directly“), which I simply had to enable in order to avoid double heads in my results.

First result after enabling „High Resolution Fix“:

(click to enlarge)

That’s the problem with news articles regarding the topic IT or AI: they become obsolete faster than you can think – I don’t mean any offense – and next time, I will go through the advanced options even when I don’t want to experiment deeper with a software.

At last, I can start generating images with my local installation of Stable Diffusion via App „Draw Things“:

(click to enlarge)

Thanks to the developer of „Draw Things: AI Generation“!

Juli 7, 2023

Bookmarks zu AI & Education

Diesen Blogbeitrag möchte ich nutzen, um einige Links aufzulisten, die ich sehr interessant finde.

VERANSTALTUNGEN ZU CHATGPT

… gab und gibt es wie Sand am Meer. An der RPTU waren wir am 20.12.22 früh dabei mit der Einladung zu einem Vortrag von Prof. Weßels unter dem Titel „ChatGPT in der modernen Lehre“. Hier der Link zur YouTube-Aufzeichnung:
https://youtu.be/_QaVNFuH6Cw

NEWS UND PUBLIKATIONEN ZU KI

Newsletter SAIL: Sensemaking AI Learning: https://buttondown.email/SAIL
„Welcome to Sensemaking, AI, and Learning (SAIL) a regular look at how AI is impacting education“
Ein spannender Überblick über die internationale Perspektive! Rückblick möglich per „full archives“.
UNESCO „ChatGPT and Artificial Intelligence in higher education: quick start guide“, 2023: https://unesdoc.unesco.org/ark:/48223/pf0000385146.locale=en
TUM „ChatGPT-4 Cookbook“:
https://www.prolehre.tum.de/fileadmin/w00btq/www/Angebote_Broschueren_Handreichungen/ChatGPT-4_Cookbook.pdf
„Unlocking the Power of Generative AI Models and Systems such as GPT-4 and ChatGPT for Higher Education: A Guide for Students and Lecturers“. University of Hohenheim, March 20, 2023:
https://digital.uni-hohenheim.de/fileadmin/einrichtungen/digital/Generative_AI_and_ChatGPT_in_Higher_Education.pdf
Ob und in welchem Maße die Verwendung von KI-Tools in deutschen Hochschulen erlaubt ist und auch wie dies kommuniziert wird, wurde kürzlich in folgendem Artikel (sicherlich eine Momentaufnahme) beschrieben:
https://hochschulforumdigitalisierung.de/de/blod/KI-Tools-und-Hochschulrichtlinien

KI-TOOLS

Last but not least hier der Link zu einem interessanten MOOC der TU Graz (Start 15.5.23):
Eine sehr gute Idee, zu zeigen, wie ein ganzer Kurs/MOOC mit Hilfe von verschiedenen KI-Tools erstellt werden kann und die Tools/Prompts, die zur Materialerstellung verwendet wurden, zu dokumentieren. Ein wenig Zeitgeschichte, wenn man später darauf zurückblickt!
https://www.tugraz.at/en/tu-graz/services/news-stories/tu-graz-news/singleview/article/ki-spricht-in-selbst-kreierte-online-kurs-ueber-sich-selbst
mit Direktlink https://imoox.at/mooc/local/landingpage/course.php?shortname=GIKI&lang=en
Auch nur einen annähernden Überblick über KI-Tools, die es wie Sand am Meer gibt (s.a. https://www.futurepedia.io/), zu bekommen, finde ich zunehmend unmöglich. Außerdem ist leider schon das Ausprobieren meist mit Registrierung und Kosten/Abos verbunden.