Interview with Garry Tiscovschi of Kreoh

‘In a demo, a use case might perform well across a few scenarios. But in real life, the system has to handle thousands of interactions, including edge cases…’


Interview with Garry Tiscovschi of Kreoh

 

This interview series brings together voices from across our GenAI Masterclass programme and the wider AI ecosystem connected to the RDI Hub. We are speaking with people who have joined us as speakers, contributors, or practitioners, and who are applying AI day to day in real organisations. Their perspectives are shaped by experience, not theory, and by what actually happens once the demos are over and the work begins.

Each interview is designed as a practical, educational piece, focused on real‑world application rather than hype. Our aim is to give SME leaders and corporate decision‑makers clear, experience‑led insight into what genuinely works with AI, where the challenges lie, and how to approach adoption in a way that is grounded, responsible, and useful.

_____________________________________________________________________________

 

Speaker bio:

Garry Tiscovschi is Co-Founder and CEO of Kreoh, a Dublin-based AI company developing applied large language model solutions for complex business workflows in tax, finance and insurance, including R&D tax claim preparation by consultants. A specialist in LLM AI engineering and a Business Post 30 under 30 honouree, Garry combines deep technical expertise with a strong background in data science and management from Trinity College Dublin. He has previously held roles at Mastercard and in education, and remains actively involved in the tech community, mentoring students and contributing to coding initiatives.

 

 

What does a successful AI deployment look like inside a mid-sized business, and why?

It starts with clear goals. You need to define the productivity uplift you’re trying to achieve, then measure performance before and after AI is introduced.

In our case, the main use case is consultants producing R&D tax claim reports. We track how many reports one consultant could complete before AI, and then after AI augmentation we look at two things: whether they’re producing more reports, and whether the quality has improved. For example, we measure whether government inquiries or red flags decrease. AI should increase both productivity and quality.

What’s often overlooked is the operational side. One approach that works well is starting with a small group of “champions” — typically your top consultants. You run a train-the-trainer model with them so they can onboard others. Five trained experts can quickly scale to train 25 or 50 people.

These champions also help configure the software around the company’s best practices and internal subject matter expertise. That ensures consistency and raises overall quality. It gives junior staff a faster start because they’re working within a system shaped by top performers.

So there are three main parts:

  • Set goals and measure impact before and after AI

  • Use a train-the-trainer model to scale adoption

  • Configure the system around top performers to maintain quality and consistency

That combination is what really makes deployment work at scale.


What separates a use case that demos well from one that survives real operations?

It comes down to one word: evals.

Evals are essentially simulation testing for AI systems. In a demo, a use case might perform well across a few scenarios. But in real life, the system has to handle thousands of interactions, including edge cases.

For example, a customer support bot won’t just handle three clean demo conversations. It will face thousands of unpredictable inputs. That’s where issues like hallucinations appear.

Strong AI deployments test across a wide range of possible scenarios, not just a few examples. They also include guardrails and safety mechanisms based on those tests.

The difference is simple: demos prove potential, evals prove reliability at scale.


How should companies think about build vs buy vs configure in 2026?

It depends on the layer you’re working on.

At the model level, most companies should not build from scratch. Competing with providers like OpenAI or Anthropic requires full focus and significant resources.
It only becomes more viable if building models is part of your core business or if you represent a truly large enterprise​.

Where most companies should focus is the application and AI harness layer – so at the software level. You can take existing models, whether open source or from major providers, and fine-tune them around your specific use case. Intercom is a good example. They fine-tuned an open source model and achieved better results for their use case than general-purpose models.

At the software layer, the approach depends on where your competitive advantage lies:

  • If the software is core to your differentiation, you should configure it heavily with a strong development team or vendor partnership

  • Your team should be trained to continuously tune and adapt it

  • In large organisations, building custom solutions may makes even more sense when even a small performance gain delivers significant value at scale

A useful way to think about it: if the software is central to your “game,” it needs to be tailored to you.


What EU AI Act or governance risks are companies walking into without realising?

One common issue is misclassifying use cases as high-risk or low-risk.

Some applications, like loan review or workplace safety monitoring, fall into high-risk categories. Others, such as AI-based emotional tracking of employees, may be inappropriate altogether. Companies need to understand where their use case sits.

Even for lower-risk cases, there are important governance steps:

  • Understand where your input data is coming from

  • Have visibility on how data flows through the system

  • Review and maintain records of the architecture, even at a high level

A key mistake is over-relying on large vendors. Companies often assume that because they’re using a major provider, compliance is automatically handled.

But responsibility sits with the organisation applying the tool. You still need to ensure:

  • The use case is appropriate

  • The data flow is compliant

  • Risks like bias or misuse are addressed

Even non-technical leaders should at least be able to review a simple data flow diagram and understand how the system works.


What does “good” look like 12 months into an AI journey?

First, you should be measuring clear before-and-after outcomes. Ideally, you can quantify productivity gains and improvements in quality.

In our case, we look at reduced government inquiries as a proxy for accuracy, and compare AI-assisted human plus AI outputs to purely human or purely AI results.

Beyond that, “good” looks like strong internal adoption. You want your initial champions or AI ambassadors actively experimenting with the system.

It’s important to allow some flexibility:

  • Some workflows should be tightly controlled

  • Others should allow experimentation

Over time, these users start developing new workflows that you didn’t anticipate. By around month three, they may already be suggesting improvements or new features.

At that point, AI becomes a collaborative process where users and providers co-develop better ways of working. That’s where further competitive advantage emerges.


Is there one AI tool you can’t live without?

I tend not to recommend a specific tool because I switch often between different open source AI harness configurations.

What matters more is the approach. I currently use an open source AI “harness” configured and deployed as our internal “Kreoh Agent”. The Kreoh Agent creates sub-agents connected to tools like Gmail, Google Drive and calendars. These agents can organise files, draft emails and carry out tasks automatically.

One thing I rely on heavily is a custom writing file. It’s a markdown file that defines how the AI should write with me. It includes:

  • Rules on what to avoid, like common AI writing quirks

  • Examples of writing styles I like

  • Examples of my own good and bad writing

It even scans for new AI writing patterns to avoid and updates itself.

That file makes a big difference. It allows consistency across whatever tools I’m using and significantly improves output quality.

 

Three key takeaways

·         AI success depends as much on operational rollout as the technology, especially using champions and structured adoption models

·         Strong AI systems are proven through large-scale “eval” testing and real-world scenarios, not just polished demos

·         Most organisations gain advantage by using existing AI models for their own workflows rather than training their own flagship models from scratch.