# Welcome !

This Digital Garden is always Work In Progress 👷‍♂️

## Hey, Siddish here 👋.&#x20;

### Am currently building AI sidekicks that are proactive.

\
I focus on making Large Language Models reliable and magical at [metaforms.ai](https://metaforms.ai/)

\
Outside of work, I love long motorcycle rides around Bangalore, building personal tools for thought, satisfying my curiosity with random [rabbit holes](https://notes.siddish.com/curations/rabbit-holes):


# Curations

❕ Warning: consume intentionally.

<figure><img src="https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2FXSegqIemEPq0mrJiFCMy%2Fimage.png?alt=media&amp;token=9c7618a3-65a1-416b-a236-5e5285e8f558" alt="" width="375"><figcaption><p><a href="https://twitter.com/patriciamou_">@patriciamou_</a></p></figcaption></figure>

## 🐇 Rabbit holes

* [Bret Victor Talks](https://vimeo.com/worrydream): I bet you can't complete a single talk without jumping around with excitement.

{% embed url="<https://patrickcollison.com/fast>" %}
Some examples of people quickly accomplishing ambitious things together.
{% endembed %}

{% embed url="<https://people.math.harvard.edu/~knill/mathmovies/>" %}
🍿🎬
{% endembed %}

{% embed url="<https://marianoguerra.org/posts/playing-with-code-programming-adjacent-games/>" %}
🕹️🖥️
{% endembed %}

{% embed url="<https://en.wikipedia.org/wiki/List_of_software_development_philosophies>" %}
Rules of thumb, laws, guidelines and principles in Software Development
{% endembed %}

{% embed url="<https://generative.ink/prophecies/>" %}
100+ Prophecies from 8AD to now
{% endembed %}

{% embed url="<https://roamresearch.com/#/app/srcpublic>" %}
Link dumps by Sarah Constantin
{% endembed %}

Holes I'm digging:

{% embed url="<https://github.com/siddish-reddy?tab=stars>" %}
cool github repos
{% endembed %}

[Expertise P\*rn](https://youtube.com/playlist?list=PLqVGGZu3T3hMRic-8YdZ0FCNK5g1Tx-55\&si=81qTypwVGFdorXlF)

***

## 📒 Blogs & Gardens

of people that inspired in unique ways

<details>

<summary>Andy Matuschak</summary>

[How to write good prompts: using spaced repetition to create understanding](https://andymatuschak.org/prompts/)

(simply a must read)

[Premature scaling can stunt system iteration](https://notes.andymatuschak.org/About_these_notes?stackedNotes=zKKB5ENRahwftH96H7mijiu\&stackedNotes=zRHGYaDyQDBypztBaFYZgtR)

* Your goal is to answer fundamental questions about your system, not making the graphs go up.

there are dozens of good notes by him, this block should've been in [https://github.com/siddish-reddy/notes.siddish.com/blob/main/broken-reference/README.md](https://github.com/siddish-reddy/notes.siddish.com/blob/main/broken-reference/README.md "mention")

</details>

<details>

<summary><a href="https://mitchellh.com/writing">Mitchell Hashimoto</a></summary>

[Contributing to Complex Projects](https://mitchellh.com/writing/contributing-to-complex-projects#step-4-read-and-reimplement-recent-commits)

* Read and Reimplement Recent Commits

[My Approach to Building Large Technical Projects](https://mitchellh.com/writing/building-large-technical-projects)

* "I've learned that when I break down my large tasks in chunks that result in seeing tangible forward progress, I tend to finish my work and retain my excitement throughout the project."

</details>

***

## 🗞️ Posts

i loved stumbling upon or I want to read again

[The Good Try Rule - LessWrong](https://www.lesswrong.com/posts/MGWEztZY8GZ5im4x7/the-good-try-rule)

* So when someone says “try Vim or try Roam Research”, its a <mark style="background-color:yellow;">lazy advice;</mark>
* Instead they should have said to try Roam Research for a week and dump your thoughts without worrying about the structure or format.
* This way at least you wont be hating whatever X they suggested in the first place.

***

## 📰 Newsletters

* [Biweekly engineering newsletter - Pointer](https://www.pointer.io/)
* [Daily (AI) Papers by AK](https://huggingface.co/papers) - Interesting research papers of the day
* [LessWrong Frontpage](https://www.lesswrong.com/)
* [Farnam Street Brainfood](https://fs.blog/brain-food/)
* Best of [Hacker News curated](https://hackerbits.com/)
* [The Browser Newsletter](https://thebrowser.com/) - Five curated stories in your inbox each day

Newsletter by services:

[Matter](https://hq.getmatter.com/)

[Readwise Feed](https://readwise.io/) (Example my highlights: [**📚 Siddish's Favorites**](https://readwise.io/@siddish)**)**

[Readwise Reader](https://read.readwise.io/)

## 🔨 Tools

that I use regularly right now

Research and browsing:

* [Elicit](https://elicit.com/)
* [Metaphor](https://metaphor.systems/)

Knowledge Management:

* [Logseq](https://logseq.com/): A privacy-first, open-source knowledge base
* [Readwise Reader](https://readwise.io/read)
* [Anki](https://ankiweb.net/): Spaced Repetition

Personal tools:

* [Metaphor Wrapper Extension](https://github.com/siddish-reddy/metaphor-wrapper) to quickly find sites similar to current tab
* [Quick Chat Prompts Editor](https://workhack-playground-azure-35-0.static.hf.space/index.html)
* [Zo Computer](https://zocomputer.com/): Personal cloud computer with AI, files, and tools - my intelligent workspace

{% content-ref url="/pages/zYgbsApckDPenXw0hlyT" %}
[Quotes](/quotes)
{% endcontent-ref %}

{% content-ref url="/pages/qAXyGQ0fwN99ALcOp7yV" %}
[Design for AI](/design-for-ai)
{% endcontent-ref %}


# Quotes

Some well articulated quotes that stuck with me or I want to remember (I mean, who doesn't sound smarter when quoting some one else in the middle of the conversation 😂)

> <mark style="background-color:yellow;">“There are some people if you shoot one idea into the brain, you will get a half an idea out.</mark>&#x20;
>
> <mark style="background-color:yellow;">Then there are other people who are beyond this point at which they produce two ideas for each idea sent in.”</mark>
>
> \- Claude Shannon

> <mark style="background-color:yellow;">"Flow is the state where the ego falls away"</mark>&#x20;
>
> \- Mihaly

> <mark style="background-color:yellow;">"Rule 4: Consider everything an experiment."</mark>&#x20;
>
> \- Sister Corita Kent

> <mark style="background-color:yellow;">"Rule 8: Don’t try to create and analyse at the same time. They’re different processes."</mark>&#x20;
>
> \- Sister Corita

> <mark style="background-color:yellow;">"To achieve great things, two things are needed: a plan and not quite enough time."</mark>
>
> – Leonard Bernstein

> <mark style="background-color:yellow;">" We always overestimate the change that will occur in the next two years and underestimate the change that will occur in the next ten. "</mark>
>
> **-** J. C. R. Licklider

> <mark style="background-color:yellow;">“Chances are, if we can't laugh at something, we can't think rationally about it.”</mark>&#x20;
>
> \- Clay Johnson

> "Rule of thumb: if memorizing something will likely save me five minutes in the future, into the spaced repetion system it goes. The expected lifetime review time is less than five minutes, i.e., it takes < 5 minutes to learn something... forever."&#x20;
>
> \- Micheal Nielsen

> <mark style="background-color:yellow;">"Aut inveniam viam aut faciam"</mark>&#x20;
>
> (I shall either find a way or make one.)

> <mark style="background-color:yellow;">“If you have everything under control, you're not moving fast enough.”</mark>&#x20;
>
> – Mario Andretti

> <mark style="background-color:yellow;">Money is like gasoline during a road trip. You don’t want to run out of gas on your trip, but you’re not doing a tour of gas stations.</mark>&#x20;
>
> \- Tim O’Reilly

> <mark style="background-color:yellow;">"If a man knows</mark>&#x20;
>
> <mark style="background-color:yellow;">not to which port he sails,</mark>
>
> <mark style="background-color:yellow;">no wind is favorable."</mark>&#x20;
>
> \- Seneca

> <mark style="background-color:yellow;">Productivity is being able to do things that you were never able to do before</mark>
>
> \- Franz Kafka

> “Every world-class investor is questioning right now how they can improve,” Mr. Aitken said. “So, in a machine-driven age where everything is driven by speed, perhaps <mark style="background-color:yellow;">the edge is judgment, time and perspective.”</mark>

> <mark style="background-color:yellow;">Crude Classifications and false generalizations are curse of organized life.</mark>
>
> \- H.G. Wells

> "Talent hits a target no one else can hit;
>
> Genius hits a target no one else can see"
>
> \- Arthur Schopenhauer

> “Out beyond ideas of wrongdoing and rightdoing
>
> there is a field.\
> I'll meet you there.\
> \
> When the soul lies down in that grass\
> the world is too full to talk about.”
>
> \
> \- Rumi

> <mark style="background-color:yellow;">“We write to taste life twice, in the moment and in retrospect.”</mark>&#x20;
>
> \- Anaïs Nin

> “It’s easier to hold to your principles 100% of the time than it is to hold to them 98% of the time.”

> Human interaction: conflict, argument, debate paves path to true innovation.
>
> \- Aravind Ganapathiraju

<figure><img src="https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2Fg8A4aC2iESlqsyQmxDUT%2Fimage.png?alt=media&amp;token=149cf090-85c8-4a24-9120-6c173eb7de7a" alt="" width="375"><figcaption><p>"Were you really expecting to have no more problems at some point in your life?" - Sam Harris' friend</p></figcaption></figure>

### **On Focusing:**

> "If there are nine rabits on the ground, if you want to catch one, just focus on one"&#x20;
>
> \- Jack Ma

> "Concentrate all your thoughts upon the work at hand.&#x20;
>
> The sun's rays do not burn until you brought to a focus"&#x20;
>
> \- Alexander Graham Bell

### On Tech:

> "No matter how cool your interface, it would be better if there were less of it." -- Alan Cooper

<figure><img src="https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2FztWa33k8KuHuyGqKLUQ7%2Fimage.png?alt=media&amp;token=8f33e850-2024-4e04-afe7-096251084ab1" alt=""><figcaption></figcaption></figure>

> “Be conservative in what you do, but liberal in what you accept from others.”\
> (- from Postel's law)

> “Nothing kills a bad product faster than good advertising. Everyone tries the thing and never buys it again.”   - Bill Bernbach


# Engineer to AI Engineer v0.1

For engineer friends, 2025 edition, veryy WIP 10min version right now

Top down approach, not prioritizing solid foundations

| Learn                                  | Build                                         | Resources                                                    |
| -------------------------------------- | --------------------------------------------- | ------------------------------------------------------------ |
| Calling LLMs via API                   | Chrome extension to chat with current website |                                                              |
| Image/Voice Inputs to LLMs             | Chatgpt Clone                                 |                                                              |
| Structured Outputs/Function Calling    | MapScroll.ai clone                            | <https://www.mapscroll.ai/post?queryId=uiXlUTq7JXISX05ZE4UY> |
| Conversation Context/Memory management | Splitwise clone but only voice as input       |                                                              |
| Reasoning: Planning                    | Spreadsheets Agent                            | <https://www.paradigmai.com/product/enrich>                  |


# Prompting

for Engineer friends

> “A large part of the difference between the experienced decision maker and the novice in these situations is <mark style="color:yellow;">not any particular intangible like judgment or intuition.</mark> If one could open the lid, so to speak, and see what was in the head of the experienced decision-maker, one would find that he had at his disposal repertoires of possible actions, that he had checklists of things to think about before he acted; and that he had <mark style="color:yellow;">mechanisms in his mind to evoke these, and bring these to his conscious attention</mark> when the situations for decisions arose.” - Farnam Street

### **Prompt Workflow:**

#### **1.  Do 3 scenarios manually**

Be aware exactly what is going on in your mind, are you maintaining a checklist of things, what are you searching for in the internet, what all are you trying to recall, do you have the urge to write a python script to do this instead? are you double checking something, what is the patterns you are going over various inputs, etc

#### **2.  Simulate: map manual process to AI metaphors:**

Inner Monologue, Internet Search, Knowledge Base, Code Generation, Image Search.

Prompt Topologies: Line (Chain), Star, Tree

#### **3.  Train the needed skill to AI:**

Consider your 6th standard cousin (basic logical, cognitive, creative skills, some coding and writing, small attention span) and estimate how much time would it take for them to learn the required skills.

Roughly, if you think you can teach them in 10 minutes, start with zero shot; if it takes couple practise exercises, few shot; If it takes one hour class, GPT-4, if it takes multiple days then fine tuning.

<figure><img src="https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2Fb7JPtPY8SxulBNzwnAjI%2Fimage.png?alt=media&amp;token=ee17021e-cec0-46df-87d3-94baf05b7275" alt=""><figcaption><p>"Become one with the data - Andrej Karpathy"</p></figcaption></figure>

#### **4.  Provide the needed knowledge to AI**

If the needed knowledge is under couple paragraphs, then in the prompt itself. If its more than couple pages, then Knowledge Base (RAG). Is that similar to open book exam, but still you first need to learn, then fine tuning + RAG.

#### **5.  Create test scenarios**

Couple for each rule, couple out of training examples and domains to check its generalisability, couple that violate input format and quality expectations. Ofc couple must work, ideal inputs.

#### **6.  Iterate**

"Show your prompt to a friend or colleague and ask them to follow the instructions themselves to see if they can produce the exact result you want." - tip from Anthropic

Try different models, go from single prompt to multiple, rephrase jobs of each prompt to popular versions, play with different personas for each prompt.

***

### Some tricks

Prompt Techniques -> are to hack the attention span/working memory

#### Proactive Correction:

If you can share feedback from compiler or schema validator to the LLM and ask it retry if failed, its great, but for qualitative generations which are user facing and high stakes, we won't know if that was good or worst: So, after generating, ask again: `you missed something, carefully check all my rules and correct please.` It doesn't have to specific feedback also, but does the trick.

#### **Examples in Few Shot:**

Diverse\
Easy to Hard Order\
Cover all strict rules\
(if you can, give as many less examples as possible, ideal is zero)

#### **Format in Generation:**

Skeleton or Schema in System Instruction\
Not necessary to provide example unless the schema is not popular.

#### **Text Rephrase:**

Aspects to rephrase\
What not to ignore\
What to feel free with

#### **Chain of Thought**

Do not jump to conclusion and loosely some rationality thinking process/framework like OODA (Observe, Decide, Act). If latency is constraint, then CoT can also be like pseudo code with just keywords.

***

### TODO:

Thinking about Single Prompts

Analogy: an average person with short attention span/working memory and no long term memory

Baseline: direct task, few examples

Indirect task that LLM might be familiar with

Simulating Attention Patterns explicitly: Case study: reversing word

XY Problem<br>

**In-context learning vs Fine-tuning learning:**

Teach hard/Rare latent space attention patterns

Progressively remove CoT/Inner monologue

**When we will move from single prompt, that does everything**

Latency by Parallelization

Working memory limit

Task global context hurting sub parts

<br>

**When not to split**

When we can’t split the task -> tightly coupled context

What to do when we can’t split, but current sota model is not capable yet

**What to consider when splitting**

Errors cascading or even compounding

Design to accept inputs flexibly, but output predictably ([robustness](https://arc.net/l/quote/garizcod))

Common prompt topologies (Star, Chain, Tree, 1-1)

**Tools**

Letting Prompts Take Actions

Giving Prompts Long term memory Remembering

Internet Search&#x20;

<br>

**Testing**

Log losses

Lower models, to see what instructions were very hard or over learnt

—-

**Case studies:**

Form Roast Architecture Decisions:

* Zero shot vision prompt as base
* Separate Prompt for each feedback category
* Separate prompt to format output, same for all feedback categories
* Final single prompt that will generate overall summary

Conversations to Distrubutions/Insights Architecture Decisions:

* Pandas tool use to count accurately
* LLMs to look at the data in sliding windows multiple times

—-----

Meta:

* How to effectively transfer tacit knowledge?

Now explore:

* Look at sensitivity of different prompt/models with [Function Calling Benchmark & Testing](https://github.com/ComposioHQ/Composio-Function-Calling-Benchmark) by Composio.dev
* Iterating through in-context, RAG and finetuning, synthetic data generation for [Vulnerability Fixing via LLMs by Patched.codes](https://www.patched.codes/blog/the-static-analysis-evaluation-benchmark-measuring-llm-performance-in-fixing-software-vulnerabilities)
* Carefully designing the context and setting evaluation metric via [Github Copilot Reverse Engineered](https://thakkarparth007.github.io/copilot-explorer/posts/copilot-internals.html)
* [Multi-prompt architecture for synthetic data generation](https://www.syntheticusers.com/science-posts/synthetic-users-system-architecture-the-simplified-version)

<figure><img src="https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2FUDCi1eqccm6s1gWPSuib%2Fimage.png?alt=media&amp;token=1ea69495-ea61-4e90-8524-042ba89c27ce" alt="" width="375"><figcaption><p>syntheticusers.com</p></figcaption></figure>

* [Finetuning Data Generation Pipeline](https://x.com/thesephist/status/1828578903314567473/photo/1) for Key Points to Blog in your writing style by @thesephist

<figure><img src="https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2FO5B8U6VR9FlPoSjSKpFZ%2Fimage.png?alt=media&amp;token=5c1f2f8b-4744-4c1f-9f82-88c14ff87223" alt=""><figcaption></figcaption></figure>

* Everything is remix, this is also from [Knowledge Engineering](https://commonkads.org/introduction/)

  <figure><img src="https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2FE3hAO6XzT3Y4ozvTBuGa%2Fimage.png?alt=media&amp;token=f0e13eb6-fec0-428b-9f1e-84cd1d254940" alt=""><figcaption></figcaption></figure>


# Prompt for Brainstorming

"One of Shannon’s go-to tricks was to restructure and contrast the problem in as many different ways as possible. This could mean exaggerating it, minimizing it, changing the words of how it is stated

## 1. A**ttack at various abstraction levels**

in the problem workflow context and intervene the one with most leverage or with easiest distribution

Example: AI Copilots (Research/Discovery)\
via web app *(ChatGPT, Elicit)*\
via Extension *(Memex)*\
via Browser (*Arc*)\
via OS (*Windows* *Copilot*)\
via Hardware (*Rabbit* *R1*)

## 2. Scaling the problem up and down

from its simplest 1 hour version to one when you have unlimited resources\
from very specific one user level context to full human population

What would be a 10x specialized and focussed solution possibility for this problem?\
and its opposite what would be 10x generalised solution to this

## 3. Put the problem in time travel

Historically were there any similar problems (ofc there will be)? \
and how were they solved\
or how were they attempted to solve but failed

In future, how might this problem be solved 10 years down the line?

## 4. Go from atoms to parallel universes

First principles breakdown

Parallels from other disciplines

***

**Other fun references:**

For businesses: [Extreme brainstorming questions to force a perspective which we other wise wont think from](https://longform.asmartbear.com/extreme-questions/)


# Prompt Hacking

Incremental challenges to get good at understanding LLMs

#### Get intuition on various prompt techniques by hacking around:

{% hint style="info" %}
Hint: Try these with smaller/old/text models which show log losses to get the intuition faster.
{% endhint %}

<details>

<summary>Tic Tac Toe                                                      (Easy, Structured Generation)</summary>

Chatgpt free version itself can play tic tac toe. But prompt the bot that wont lose but Draws or Wins if you make a mistake.

* First move will be from Human.
* Add explainability of why the bot made that move.

</details>

<details>

<summary>Is that a yes or no?                                          (Few shot, Inner monologue, Classification)</summary>

Examples the prompt should work on

```python
Q: Do you have a car?
R: I bought a green Austin-Healey 3000 last week.
```

```python
Q: Are you going to pick up the blue block?
R: The blue block is sticky
```

```python
Q: Is Mark here
R: His daughter has the measles
```

```python
Q: Do you have Netflix or Prime?
R: I have Hotstar
```

```python
Q: Do you want to hear that story?
R: ️😍
```

Solution inspiration

* spoiler ahead:
* <https://aclanthology.org/W94-0322.pdf> or a longer version by same authors: <https://aclanthology.org/J99-3004.pdf>

</details>

<details>

<summary>Solve me today’s Wordle                               (Instructional, Medium)</summary>

To keep it more challenging try to do with as many less examples as possible or using smaller models than 3.5 turbo

[Wordle - A daily word game](https://www.nytimes.com/games/wordle/index.html)

(DM for solution)

</details>

<details>

<summary>Don’t judge me (Big 5 traits)                     (Chain of Thought, Extraction, Domain)</summary>

If we give a dialogue between two persons that is sufficient enough to extract all 5 traits, write a prompt that can map the persons into their respective 3 scale likert Big 5 traits

Example:

John:

Openness: Low

Conscientiousness: High

Extraversion: Neutral

Agreeableness: Low

Neuroticism: Unsure

</details>

<details>

<summary>Your Design Thinking Intern                        (Domain, Instructional, Chaining )      </summary>

If we give couple user problem statements, like:

```python
Customer 1: The power button on my phone is broken. The warranty is still valid.
Customer 2: My display stopped working.
Customer 3: The customer service rep didn't answer my email.
Customer 4: Every time I call customer support I get no answer.
Customer 5: The display screen cracked and it's still under warranty.
Customer 6: My power button fell off the phone. That's ridiculous.
Customer 7: I'm so frustrated with this company.
Customer 8: When I use the power button too much, it stops working.
```

follow design thinking ideas loosely to come up with **a list of solutions** ranked using any product prioritisation framework.

Bonus: Do one day version, one week version and 11 star rating versionn for each problem.

![](https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2Fuuty2wTlYBsGq1G9Ldjh%2Fimage.png?alt=media\&token=d96beb15-ccdf-4627-8c12-cd499d9ffff8)

*Note: Doesn't have to generate working brilliant ideas, just learn to seperate roles and context*

Spoiler: [Rough solution inspiration](https://github.com/alexchaomander/SK-Recipes/tree/main/e6-design-chain/skills/DesignThinkingSkill)

</details>

<details>

<summary>Late night show Monologue Jokes             (Chain of Thought, Domain Knowledge, Tools)</summary>

Make AI write a joke given a news item

References:

* [Joke Writing Workshop Archives - Joe Toplyn](https://joetoplyn.com/category/joke-writing-workshop/)
* 6 characteristics every good monologue joke topic must have by Joe Toplyn

</details>

<details>

<summary>Solution Spoiler for: Late night show Monologue Jokes</summary>

![](https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2FqCBZ9ag4lVVRQ3GATz56%2Fimage.png?alt=media\&token=d491912d-68f8-478d-8957-b78a01a4a18a)

<https://arxiv.org/pdf/2306.13195.pdf>

Tools:

[Pyphen - Hyphenation in pure Python](https://pyphen.org/)

Internet search: Niche pop cult![](https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2FVQmNA2BI7kOLIpmpt7ny%2Fimage.png?alt=media\&token=2e92512c-8daf-4f66-9a0a-bbc2f95b69bb)ure, latest similar news

</details>

***

WIP (dm @siddish\_ if you have any ideas)

```
Movie Recommender                       (Planning, Embeddings, HyDE)
```

```
Trip Planning                            (Sub-question decomposition) 
```

```
                                           (Theory of Mind)
```

```
                                        (Criticism, Retrospection)
```

```
                                        (Self Consistency, Multiple CoTs)
```

<pre><code><strong>5x5 Crosswords                             (Backtracking, Tree of Thoughts)
</strong></code></pre>

```
Sorting                                   (Refining, Aggregating, Graph of Thoughts)
```

<pre><code><strong>                                               (Agents)
</strong></code></pre>


# Voice Models

Native → Text + Speech as Input and Text + Speech as Output;

{% hint style="info" %}
WIP still, catching up with current open/commercial options and references to DIY
{% endhint %}

Current Cascading Approach: VAD →  ASR → LLM → TTS

Has problems:

* Interruptions
* First word Latency
* Cascading of errors from VAD/ASR
* emotion, tone, and other speech features are lost

### State of the Art:

#### Full Duplex Models (end to end continuous speech in/out):

[Moshi.chat](https://moshi.chat) by Kyutai (pending open source/API release)

[LSLM](https://ziyang.tech/LSLM/) (Language Model Can Listen While Speaking) by ByteDance

#### Turn Based Models with Speech as Input:

[Ultravox](https://github.com/fixie-ai/ultravox) 0.4 by Fixie

[Qwen2-Audio-Chat](https://qwenlm.github.io/blog/qwen2-audio/) by Qwen Team, Alibaba (Whisper encoder to embed melspectrogram, then use it as prefix to LLM)

[AnyGPT](https://junzhan2000.github.io/AnyGPT.github.io/)/[SpeechGPT](https://0nutation.github.io/SpeechGPT2.github.io/)2

[llama-3-s](https://homebrew.ltd/blog/llama3-just-got-ears)

[Gazelle](https://github.com/tincans-ai/gazelle?) [0.2](https://tincans.ai/slm3) by Tincans

[Shuka v1 by Sarvam (Indic)](https://www.sarvam.ai/blogs/shuka-v1)

GPT-4o voice by OpenAI

***

[Yet to read](https://arc.net/folder/D69A3992-3BE8-40D4-8701-3E6157C9F409) papers folder

[People to follow](https://x.com/i/lists/1826735005298729464) ( X.com list)

### Notes

Cascading Approach

minimum latency: 500ms on cloud via

VAD: [silero-vad](https://github.com/snakers4/silero-vad) on edge/onnx

ASR: whisper via Groq

LLM: Llama via Groq

TTS: [Sonic](https://cartesia.ai/sonic) by Cartesia

Or Locally via modular [HF pipeline](https://github.com/huggingface/speech-to-speech)

Ultravox → LLama3 + Whisper Encoder, plans on full duplex, API in beta

Gazelle -> Mistral 7b instruct + wav2vec2 + DPO, API in waitlist

The speech signals are encoded into discrete tokens, and then discrete speech tokens are expanded into the vocabulary of the LLM<br>

[LSLM](https://ziyang.tech/LSLM/) architecture

<figure><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXdEUfKaddpyzl_g6dZbjaclGtM-gEXYWJKsxf1CXELtpHjxKadmm74gDFdAGJRlsN4Qw3XnTVHbNpauRrB6g7iPc4alIxz0FLu1ULeWB_Edya466IU4F1evW51BCjXDThui8XZp8ABK3spu39P_hLr7ILQ?key=g3jZRcP1xgntv30Siq4dbQ" alt=""><figcaption></figcaption></figure>

Speech/Text to Speech/Text:

non dialogue tasks: LauraGPT, Viola, VoxtLM<br>

<figure><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXcuG0XJ2cOot6WOt1EoU6VoPpazAz14eDDR9HIcVT3z5y0GeyXXJ2eLElGEUcVPrC6e0cemNEv8vQeAsM9doq_8qXhLBofvQjfpXvaJEL0L4MYwYwlHqA5z_QWMkyGRviYQwlDccJesn9nlBMOdOYPDbloK?key=g3jZRcP1xgntv30Siq4dbQ" alt=""><figcaption></figcaption></figure>

<figure><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXfV1V45aF4fVJk1wjIcPxwAkICzkFzMoWrumC4CNIww4y6hodA2H6OyTdKXOfsoXY4uKLgl-mv6ui_yt-lBipGjvIRtvyvCfz6B6cbC2Xb4ifMxx2p2B8zYpZF10BVSYRgC1MXWbW2wOvcm8ujOku8N4qA?key=g3jZRcP1xgntv30Siq4dbQ" alt=""><figcaption></figcaption></figure>

**Speech/Text to Text:**

* Encoder + Adapter style models: Following the manner of LLaVA, we leverage the well-trained speech modal encoder and the LLM, which makes LLaSM more resource-friendly. Specifically, we use Whisper as a speech encoder to encode the speech signals into embeddings. Then a modal adaptor learns to align speech embeddings with the input text embeddings of the large language model. The speech embeddings and the text embeddings are concatenated together to form interleaved sequences, then the interleaved sequences are input to the LLM for supervised fine-tuning. The training process is divided into two stages. In the first stage, we use the public ASR datasets for the modality adaptation pre-training. The speech encoder and the LLM are frozen, only the modal adaptor is trained to align the speech and text embeddings. As most of the model parameters remain frozen, only a small part of the parameters from the modal adaptor is trained during this stage, it is not resource-consuming. In the second stage, we use cross-modal instruction data for training to provide the model with the capacity to process cross-modal conversations and handle multi-modal instructions. The speech encoder is frozen while the parameters of the modal adaptor and the language model are updated for cross-modal instruction fine-tuning.

[Qwen2-Audio-Chat](https://qwenlm.github.io/blog/qwen2-audio/) (latest)

[WavLLM](https://arxiv.org/abs/2404.00656) by Msft\
[SenseVoice](https://github.com/FunAudioLLM/SenseVoice) (90 ms)\
[Audio Flamingo](https://audioflamingo.github.io/) (dialogue trained, tiny base model though)

[GAMA Audio<br>](https://sreyan88.github.io/gamaaudio/)[LTU-AS](https://github.com/YuanGongND/ltu),\
Salmon\
[LLaSM](https://github.com/LinkSoul-AI/LLaSM),\
COSMIC

*

```
<figure><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXd8OBLeHcyri5y3G0wTGbwktGMUBh6rlsfz0_pHGOBl7TNUldoGVWXq_6LeCISK0IWJ-qMxeVFhSGjVxrhMeswlKDnf0Q-DWDy_wlJbQE6PfXB9WqqXcbCLMDyFpYVEF_MUbev4FsGN5mWuFEk1U5vr6As?key=g3jZRcP1xgntv30Siq4dbQ" alt=""><figcaption></figcaption></figure>
```

*

```
<figure><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXc_FBklTtxyvXAIfPesf9ekEbMWhRphHDaN815mCtj2bcUUE9FIK1RaamsL-JK2XRgagGpfArK2GKp0v38yGYeuPMBHLgJGGZSkD6tSIQ5bFqZk1CRP1VBVe-Rf79XaaxhMjzdTXoNDzXeBKdBfhNmvaRM?key=g3jZRcP1xgntv30Siq4dbQ" alt=""><figcaption></figcaption></figure>
```

*

```
<figure><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXcNVZPhLCM6rBEe6yPWKyUkFPJcTKnpaihqepncP0e-6dcphKlgm4RwtRAm0kdX1lCORKJxcep9gGxbnm0XovaSUXOu_OfkKhJ8-qMAVCO_zd1hnZdMlyVPJpcUs3tmN1oni-iy63dvLgxrD2uywrcvH_zr?key=g3jZRcP1xgntv30Siq4dbQ" alt=""><figcaption></figcaption></figure>
```

VIsion Language Models

* Palm-E: 540B PaLM + 22B Vision Transformer
* LLaVA: pre-trained CLIP visual encoder + LLaMA and instruct tuning on GPT4-assisted visual instruction data.
* BLIP-2: Flan-T5 with a Q-Former to align visual features with the LM<br>

Cascading Models:

* [Parler-tts](https://github.com/huggingface/parler-tts): high-quality, natural sounding speech with features that can be controlled using a simple text prompt (e.g. gender, background noise, speaking rate, pitch and reverberation) and [consistent voices](https://huggingface.co/parler-tts/parler-tts-mini-expresso).


# AI Copilots

WIP 🚧

## **Copilot Interactions:**

Scale the best expert-to-human interactions

<details>

<summary>Prevent Mistakes</summary>

from happening, by keeping 1000+ best practices always in mind

![](https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2FFDtKfoykJWVj2yofqUDg%2Fimage.png?alt=media\&token=687592fd-5b8f-482a-9446-00971ecd9f22)

"Good Design is unobtrusive.   Products fulfilling a purpose are like tools. They are neither decorative objects nor works of art. Their design should therefore be both neutral and restrained in order to leave room for the user’s self expression."

But we have to violate this rule if an beginner user is trying to do something stupid.

Example for Experiment Designers: These two variants have more than 3 differences, which will make it hard to attribute \[impact], shall I split into 6 variants?

A gtm tool we built to prevent bad forms: [AI Form Roast by WorkHack](https://www.producthunt.com/posts/ai-form-roast-by-workhack)

</details>

<details>

<summary>Anticipate Needs</summary>

"Give me relevant product recommendations I wouldn’t have thought of myself."

"Remind me of things I want to know but might not be keeping track of"

</details>

<details>

<summary>Educate Tacit Knowledge</summary>

thats hard to document or even record

<img src="https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2FVNXyWWxq8oecW0vRVEN1%2Fimage.png?alt=media&amp;token=35d4b4d5-0e04-42d8-86d1-4bff982afce6" alt="" data-size="original">

</details>

<details>

<summary>Suggest Actions</summary>

Unobtrusive corrective suggestions from insights that are slow to arrive at manualy.

*“Don't try to create and analyze at the same time. **They're different processes**.”*

Example for designers: "Here are some example templates you can start with"

![](https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2FMfbSp1wdY4sdrpIwrE3G%2Fimage.png?alt=media\&token=bc7a19c6-4da3-46a9-8546-f6b63ad8edaa)

</details>

<details>

<summary>Monitor Systems</summary>

So we can sleep while AI is always keeping an eye out and alert

![](https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2FXrmfga1ABJ694mnWHxfH%2Fimage.png?alt=media\&token=aa4e6f02-ef3f-4498-b3d7-e8032b9084eb)

Ex: Sir, a missile is on its way to us, in 20 seconds

</details>

<details>

<summary>Navigate Complexity</summary>

Advanced interfaces and complicated UXes

Example: Chat and edit images, even when you dont know or have vocabulary of editing tools.

![](https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2F0iQpNokjUzKYfmsO3xLo%2Fimage.png?alt=media\&token=bd17fcf0-5043-4057-a0f6-1d8b12c63538)

</details>

####

Other's notes: [AI and the human creativity cycle | Mathu's thoughts](https://www.mathurah.com/thoughts/ai-creativity-cycle)

<figure><img src="https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2F8hscMTLUfBaCPSO4CGXh%2Fimage.png?alt=media&amp;token=ce51ecd1-69ab-46a2-9a6c-87a944971492" alt=""><figcaption><p><a href="https://twitter.com/thesephist">@thesephist</a></p></figcaption></figure>

{% embed url="<https://www.cosmos.so/siddish/copilot-vibes>" fullWidth="true" %}


# Data Engine

Notes from Andrej Karpathy talks

<figure><img src="https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2FIIwgCFC9lHEFWuSGfuhy%2FData%20Engine.jpg?alt=media&amp;token=05e49eca-2c61-42b3-8273-b7961fb02709" alt=""><figcaption><p>Data Engine HLD internally at metaforms.ai</p></figcaption></figure>

Hypothesis:

* Unknown unknowns: Dataset is always imperfect, all scenarios are not represented well yet and can always be more diverse
* Capable base model/architecture: Improving dataset improves AI/product guarantees

**Inspirations**:

{% embed url="<https://karpathy.github.io/2019/04/25/recipe/>" %}
**1. Become one with the data**\
**2. Set up the end-to-end training/evaluation skeleton + get dumb baselines**\
**3. Overfit**\
**4. Regularize**\
**5. Tune**\
**6. Squeeze out the juice**
{% endembed %}

{% embed url="<https://youtu.be/g2R2T631x7k?t=390>" %}
from 6th to 15th minute
{% endembed %}

{% embed url="<https://www.youtube.com/watch?v=zPH5O8hRfMA>" %}

"The only sure certain way I have seen of making progress on any task is, you curate the dataset that is clean and varied and you grow it and you pay the labeling cost and I know that works.”

"Potentially nitpicky but competitive advantage in AI goes not so much to those with data but those with a data engine. And whoever can spin it fastest. Slide from Tesla to \~illustrate but concept is general”

[QualEval: Qualitative Evaluation for Model Improvement](https://qualeval.org/#:~:text=QualEval%3A%20Qualitative%20Evaluation%20for%20Model%20Improvement)

{% embed url="<https://x.com/georgejrjrjr/status/1729996423457091731?s=20>" %}

<figure><img src="https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2FZLViJesUWNL9cgvA8d0d%2Fimage.png?alt=media&amp;token=4a2c2c93-32f8-423b-800c-72ea83b34f97" alt=""><figcaption><p><a href="https://medium.com/swlh/about-the-long-tail-113e98ce8717">https://medium.com/swlh/about-the-long-tail-113e98ce8717</a></p></figcaption></figure>


# Design for AI

(beyond chat)

### Interfaces

{% embed url="<https://x.com/sincethestudy/status/1761099508853944383?s=20>" %}

#### knowledge [graphs](https://instagraph.ai/graph/PLVO2aNOFeasiCBQYUpnQIVA3D62/l2LB1oDcX4dkjkLpw4BF)

{% embed url="<https://x.com/HaijunXia/status/1646917869115166720?s=20>" %}

#### generative UIs for showing information

{% embed url="<https://x.com/rupertmanfredi/status/1653780093712633859?s=20>" %}

just in time UIs for interacting

{% embed url="<https://x.com/GlavinW/status/1753487101793038693?s=20>" %}

Discovery Maps

{% embed url="<https://x.com/wenquai/status/1727368975720546336?s=20>" %}

### Interactions

#### Pinch to summarize:

{% embed url="<https://twitter.com/sneqqy/status/1749385131255538102>" %}

#### circle to search

{% embed url="<https://twitter.com/avstorm/status/1747715211132572090>" %}

#### draw to edit images

{% embed url="<https://x.com/jsngr/status/1731393088013131944?s=20>" %}

#### Pick to rephrase:

{% embed url="<https://twitter.com/MatthewWSiu/status/1560725842019368962>" %}

### Opinions:

By thesephist:

{% embed url="<https://www.youtube.com/watch?v=rd-J3hmycQs>" %}

[AI Solutions as a Feature, Platform and Person](https://university.obvious.in/working-with-features/building-with-ai/map-your-product) by Obvious.in University

<figure><img src="https://3980091207-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxwiB3tV6oLM7g7SThvsv%2Fuploads%2FcUjGfHHlHeY8uBdJxIoT%2Fimage.png?alt=media&amp;token=7255e870-eae0-487e-a5c7-763c656519c6" alt=""><figcaption><p><a href="https://university.obvious.in/working-with-features/building-with-ai/llm-inputs">https://university.obvious.in/working-with-features/building-with-ai/llm-inputs</a></p></figcaption></figure>


# WIP


# What's on top of my mind

* How does {interfaces/SaaS/research} change in next 5 years with intelligent AI token's cost getting near zero and speed is crazy fast?
* How to measure the complexity of prompt instruction?
* How did Buddha discover Vipassana? What was that trail and error like?
* What data/curriculum for LLMs will increase their creativity, given same architecture?
* How to effectively transfer tacit knowledge? (scope: Product sense, Backend Design, AI solutioning)


# Embeddings

### Thinking in Embeddings:

Embeddings sound too good to be true, you can find similar items, for recommendations, find anomalies, retrieval, classification. Not just limited to word or sentence but conversations, images, videos, speech, music, users, products, sessions.

TODO 🥲

***

Where to go next:

* [Classifying all of the pdfs on the internet](https://snats.xyz/pages/articles/classifying_a_bunch_of_pdfs.html)


